The review
- Asked does this read as a clamp
- This push yes, and the suite is green
Topic · Looks right
Gist
The patch reads as if a person wrote it. Proof fails the merge when an approved shall has no witness on this branch. CodeRabbit still reviews the diff.
proof audit --fail-level warn
The generic verify question lives on verify agent code. Trust before merge names Claude and Copilot. This page is the looks-right miss.
01 · The look
A coherent helper can still implement a cap nobody sold. The review asked whether the code looked like code. It did.
The verify question lives on verify code an AI agent wrote. That page is the circular proof: the model wrote the function and the test that agrees with it. This page is earlier in the review. A person (or a PR bot) already read the patch. It looked fine. The approved shall was never a comment on the diff.
Two jobs sit in the same complaint. proof audit --check tests_pass still owns running the suite the agent also wrote. If that check is red, you already have a failure. The looks-right problem is the other job: the suite is green, the review is green, and the 40% cap was never a condition.
func ClampDiscount(pct int) int {
if pct < 0 {
return 0
}
if pct > 100 {
return 100
}
return pct
}
That helper is the kind of thing a model writes well. Tests at 0, 50, and 100 pass. A reviewer nods. The sold rule was never more than 40%. Nothing in the file asked that question.
02 · The exhibit
The review can stay green while the shall is unwitnessed. Click the tabs.
The review
The shall
Never more than 40%. Not a branch. Not a comment.
Not in the diffThe review
Still green. Still a clamp to 100.
Still looks rightProof
Same helper. Two questions. Click the tabs.
| Who | What they notice | What they lose |
|---|---|---|
| Human review | Whether the patch reads as code | The sold cap was not on the diff. Plausible is the miss. |
| CodeRabbit / the PR bot | Style, likely bugs, the comment it would leave | It still reviews the diff. It does not hold the shall. That split lives on Proof vs CodeRabbit. |
| The agent's own tests | The inputs it already knew | Agreement with the function is cheap when both came from the same window. That loop lives on verify agent code. |
| Proof | The shall still has a witness, or the merge stays red | Proof does not know what you meant if nobody signed it. Jama still authors the programme. |
We have not run Proof against CodeRabbit, Copilot, and a human review on a frozen corpus of plausible-wrong patches. The loss is named, not scored. Keep the review. Keep the PR bot. Neither is the merge gate for an approved shall.
03 · The honest loss
If nobody signed the 40% cap, there is nothing for proof audit --fail-level warn to fail on except the suite you already have.
Jama still wins at programme authoring. CodeRabbit still reviews the diff. Copilot still writes the helper. FRETish is 288 templates, not free English. 100% MC/DC on the written decision still misses a partition that was never a condition. That instrument stays on MC/DC for Go.
Catching a silent break of a shall you already hold lives on when an agent silently breaks a requirement. Trust before merge names the tools on trust Claude or Copilot before merging. Who signed the intent stays on who verifies the intent.
04 · Nearby questions
How do I verify code that an AI agent wrote is actually correct? The circular proof. Verify agent code.
How can I trust code generated by Claude or Copilot before merging it? Named tools. Trust before merge.
AI generates our specs and our code. Who verifies the intent is right? Specs and code from the same model. Who verifies intent.
What's a reliable way to gate AI pull requests on correctness? The gate argument lives on Proof vs CodeRabbit. Do not mint a twin.