For coding agents
More work you can safely hand to agents.
Coding agents propose the change. Proof keeps the agreement about what must stay true and brings back the evidence to accept it.
The live jsonparser graph answers anonymous MCP clients. No account, no token.
01 · Evidence
Start from a failure you can run.
The agent reproduces the issue. It does not guess whether a vague report sounds plausible.
A bug report names a symptom. A reproducer is a test that fails until the fix lands. The agent starts there.
The finding names what closing it takes. The agent cannot swap in an easier test. A feature, a refactor, or a migration starts the same way, from evidence agreed before the code is written.
- Kind
- known issue · fixed on master
- Violated requirement
- SYS-REQ-009
- Reproducer
- a test that pins the break, in your repository
- Observed
- unrelated structure is corrupted
- Required to close
- the reproducer passes, and the obligations on the requirement run again
The agent starts from evidence it can run.
02 · Read
Navigate by intent, not files.
Code tells the agent what the system does. Proof tells it what must stay true.
An agent that reads only the repository has to guess the promise from the code. The guess never appears in the diff.
Before it writes a line, the agent reads what this component promises, why the behavior exists, and what depends on it.
03 · Reach
Know the blast radius before editing.
The agent asks what a change can reach. It gets the answer before the first edit.
The answer is not a file list. It is the intent the change can touch, the hazards on that intent, and what already went wrong there.
An edit that looks local can break a parent requirement. That shows up before the work starts, not after the release.
Returned to the agent
Relevant intent
SYS-REQ-009 · SYS-REQ-069 · SYS-REQ-110
Blast radius
1 component · 3 files · 8 tests · 3 hazard obligations
History
1 related defect record · 2 defect classes
04 · Fix
Make the fix. Proof grades it.
The agent does not close the work with a summary. It produces what the requirement asked for.
A reproducer checks the reported instance. Wider closure needs a defined failure class, checks for sibling cases and evidence for the relevant hazards. Running the same test again does not establish that broader claim.
For a feature or refactor, connect each affected promise to the checks and observations that support it. Keep missing mappings and untested behavior visible. A review or demonstration has a different scope from an executed automated check.
When a recorded evidence basis changes, the applicable freshness checks can mark the result stale. Unrecorded dependencies can be missed. Renew the affected evidence and inspect its results before deciding whether the change is ready.
05 · Retained
Leave the evidence behind.
After the session ends, the evidence stays on the requirement.
A fixed issue leaves a defect record in the repository. A feature, a refactor, or a migration leaves a change record. The next session opens it as context. The card here is one of those records.
Evidence stays tied to the code version and the graph version that produced it. It is verified for that revision, stale and waiting, or still missing. A passing check supports its stated claim. Package readiness also depends on required evidence and policy; acceptance is a separate authorized decision.
- Type
- feature
- Title
- v1.5.0: Config, name aliases, streaming ReaderParser.
- Requirements
- SYS-REQ-115, SYS-REQ-116
- Impact review
- 19 requirements re-read across 4 files, every row fingerprinted
- Owner
- human:buger, target release v1.5.0
06 · The assurance model
You decide what still needs a person.
An agent can propose a requirement. Out of the box, it can approve one too. You choose which levels stay open.
Proof ships open. A new project starts with agent_autonomous_for: all: true. An agent can approve at any assurance level, and the record names the agent and the level. Human-only requirement approval must be configured. Set all: false and list the levels you delegate under assurance_levels. An empty list reserves every level for a named person.
The threshold is the assurance level, A to E. The principle is from NPR 7150.2. The higher the consequence, the more evidence and judgment the change needs. A is human safety. C is recoverable production infrastructure. E is demos. You set the level per component and per spec. Any single requirement can override it.
Until you narrow the policy, an agent may approve at A as readily as at E. Narrow it, and the levels split. Delegate C, D, and E, and keep A and B. An approval at A then waits for a person who owns the code. An approval at C records the agent and the level. proof approve refuses a level you have reserved, and the refusal names the setting. The record says who approved, at which level, and under which policy.
You decide which levels need your judgment. That decision lives in one place. Where the consequence is real, reserve the level, and a person approves it by name. Widen what the agent may approve from there, by scope and by consequence, as the record earns it. The agent that writes the change cannot quietly redefine what counts as done. Editing an approved requirement marks its approval stale. Any new approval is on the record, with the actor and the level.
A person validates every finding before it reaches you. That check sits outside this setting.
An agent may
- propose a requirement, as a draft
- approve at any level your policy leaves open, including every level until you narrow it, recorded as the agent
- attach evidence to an obligation
- run the obligations again
- open a known issue
Only a person may
- decide which levels stay open to an agent. That setting lives in your repository and moves in a commit you review
- approve at the levels you reserve
- decide which misses from public work are published at all
Recorded Proof operations carry actor information in the activity log. This is not a complete recording of everything an agent does in external tools. The log lives in your repository, next to the evidence, and you can read it without us. Actions in the left column run inside your repository, not through the hosted endpoint. Token scope, repository boundaries, and permissions are on the trust page.
Who checks the checker → What agents may read, propose, and approve →
07 · Your agent
Use the agent you already have.
Proof does not ship a coding agent. It does not ask you to replace the one you use.
The agent keeps its own workflow, prompts, and habits. Nothing in your repository has to be rewritten for a particular agent.
claude codecursoryour own tooling
Proof does not require a particular agent.
Bring any coding agent. Proof keeps the agreement about what must stay true.
08 · Try it
Try it now. No signup. No token.
Point your coding agent at the public jsonparser graph. Nothing is installed. Nothing is registered.
Through Proof MCP
- read requirements, with the obligations and hazards on them
- read what a requirement depends on, and what depends on it
- read known issues and findings
- read reproducers, and the tests bound to a requirement
- read change records and their history
- calculate the blast radius of a proposed change
- read and search the audited source that evidence references
- post a comment on a finding, only where the token allows it
Inside the customer repository
- propose requirement changes
- attach evidence
- run verification
- open issues
- create change records
- approve where policy allows
MCP gives the agent the model. The agent acts through the customer’s repository, CI, and permissions.
Every read carries the graph version it was answered against. Two runs a week apart can be compared.
Tool names and permission scopes are settled when Proof is installed, because they follow the shape of the repository.
The whole graph, over MCP → The methodology behind the evidence →
See it run against a real repository.
One component. One change. The answer the agent gets before it edits anything.
seeded product demo; read the labels first
Request demo →