For coding agents

More work you can safely hand to agents.

Coding agents propose the change. Proof keeps the agreement about what must stay true and brings back the evidence to accept it.

The live jsonparser graph answers anonymous MCP clients. No account, no token.

01 · Evidence

Start from a failure you can run.

The agent reproduces the issue. It does not guess whether a vague report sounds plausible.

A bug report names a symptom. A reproducer is a test that fails until the fix lands. The agent starts there.

The finding names what closing it takes. The agent cannot swap in an easier test. A feature, a refactor, or a migration starts the same way, from evidence agreed before the code is written.

KNOWN ISSUE Was violated
Kind
known issue · fixed on master
Violated requirement
SYS-REQ-009
Reproducer
a test that pins the break, in your repository
Observed
unrelated structure is corrupted
Required to close
the reproducer passes, and the obligations on the requirement run again
Drawn to shape The record is real, and fixed: KI-3 on buger/jsonparser, closed by DEFECT-260726-MFPA. See the same record.

The agent starts from evidence it can run.

02 · Read

Navigate by intent, not files.

Code tells the agent what the system does. Proof tells it what must stay true.

An agent that reads only the repository has to guess the promise from the code. The guess never appears in the diff.

Before it writes a line, the agent reads what this component promises, why the behavior exists, and what depends on it.

PROJECT COMPONENT REQUIREMENT what it promises depends on hazards obligations code and documentation evidence
Fig. 01 · The agent walks intent. The call stack is only part of it.

03 · Reach

Know the blast radius before editing.

The agent asks what a change can reach. It gets the answer before the first edit.

The answer is not a file list. It is the intent the change can touch, the hazards on that intent, and what already went wrong there.

An edit that looks local can break a parent requirement. That shows up before the work starts, not after the release.

Returned to the agent

Relevant intent

SYS-REQ-009 · SYS-REQ-069 · SYS-REQ-110

Blast radius

1 component · 3 files · 8 tests · 3 hazard obligations

History

1 related defect record · 2 defect classes

Answered from the graph, for the question I need to change Set() for nested arrays. What can this affect?

04 · Fix

Make the fix. Proof grades it.

The agent does not close the work with a summary. It produces what the requirement asked for.

A reproducer checks the reported instance. Wider closure needs a defined failure class, checks for sibling cases and evidence for the relevant hazards. Running the same test again does not establish that broader claim.

For a feature or refactor, connect each affected promise to the checks and observations that support it. Keep missing mappings and untested behavior visible. A review or demonstration has a different scope from an executed automated check.

When a recorded evidence basis changes, the applicable freshness checks can mark the result stale. Unrecorded dependencies can be missed. Renew the affected evidence and inspect its results before deciding whether the change is ready.

the agent changes CODE * REQUIREMENT ? DOCUMENTATION ? EVIDENCE STALE review or renew evidence
Fig. 02 · Confidence is withdrawn. The evidence stays.

05 · Retained

Leave the evidence behind.

After the session ends, the evidence stays on the requirement.

A fixed issue leaves a defect record in the repository. A feature, a refactor, or a migration leaves a change record. The next session opens it as context. The card here is one of those records.

Evidence stays tied to the code version and the graph version that produced it. It is verified for that revision, stale and waiting, or still missing. A passing check supports its stated claim. Package readiness also depends on required evidence and policy; acceptance is a separate authorized decision.

CHG-260728-H6ER Active
Type
feature
Title
v1.5.0: Config, name aliases, streaming ReaderParser.
Requirements
SYS-REQ-115, SYS-REQ-116
Impact review
19 requirements re-read across 4 files, every row fingerprinted
Owner
human:buger, target release v1.5.0
Public audit evidence Read from proof/changes/CHG-260728-H6ER.yaml on jsonparser master. See the same record.

06 · The assurance model

You decide what still needs a person.

An agent can propose a requirement. Out of the box, it can approve one too. You choose which levels stay open.

Proof ships open. A new project starts with agent_autonomous_for: all: true. An agent can approve at any assurance level, and the record names the agent and the level. Human-only requirement approval must be configured. Set all: false and list the levels you delegate under assurance_levels. An empty list reserves every level for a named person.

The threshold is the assurance level, A to E. The principle is from NPR 7150.2. The higher the consequence, the more evidence and judgment the change needs. A is human safety. C is recoverable production infrastructure. E is demos. You set the level per component and per spec. Any single requirement can override it.

Until you narrow the policy, an agent may approve at A as readily as at E. Narrow it, and the levels split. Delegate C, D, and E, and keep A and B. An approval at A then waits for a person who owns the code. An approval at C records the agent and the level. proof approve refuses a level you have reserved, and the refusal names the setting. The record says who approved, at which level, and under which policy.

You decide which levels need your judgment. That decision lives in one place. Where the consequence is real, reserve the level, and a person approves it by name. Widen what the agent may approve from there, by scope and by consequence, as the record earns it. The agent that writes the change cannot quietly redefine what counts as done. Editing an approved requirement marks its approval stale. Any new approval is on the record, with the actor and the level.

A person validates every finding before it reaches you. That check sits outside this setting.

An agent may

  • propose a requirement, as a draft
  • approve at any level your policy leaves open, including every level until you narrow it, recorded as the agent
  • attach evidence to an obligation
  • run the obligations again
  • open a known issue

Only a person may

  • decide which levels stay open to an agent. That setting lives in your repository and moves in a commit you review
  • approve at the levels you reserve
  • decide which misses from public work are published at all

Recorded Proof operations carry actor information in the activity log. This is not a complete recording of everything an agent does in external tools. The log lives in your repository, next to the evidence, and you can read it without us. Actions in the left column run inside your repository, not through the hosted endpoint. Token scope, repository boundaries, and permissions are on the trust page.

07 · Your agent

Use the agent you already have.

Proof does not ship a coding agent. It does not ask you to replace the one you use.

The agent keeps its own workflow, prompts, and habits. Nothing in your repository has to be rewritten for a particular agent.

claude codecursoryour own tooling

Proof does not require a particular agent.

Bring any coding agent. Proof keeps the agreement about what must stay true.

08 · Try it

Try it now. No signup. No token.

Point your coding agent at the public jsonparser graph. Nothing is installed. Nothing is registered.

Through Proof MCP

  • read requirements, with the obligations and hazards on them
  • read what a requirement depends on, and what depends on it
  • read known issues and findings
  • read reproducers, and the tests bound to a requirement
  • read change records and their history
  • calculate the blast radius of a proposed change
  • read and search the audited source that evidence references
  • post a comment on a finding, only where the token allows it

Inside the customer repository

  • propose requirement changes
  • attach evidence
  • run verification
  • open issues
  • create change records
  • approve where policy allows

MCP gives the agent the model. The agent acts through the customer’s repository, CI, and permissions.

Every read carries the graph version it was answered against. Two runs a week apart can be compared.

Tool names and permission scopes are settled when Proof is installed, because they follow the shape of the repository.

See it run against a real repository.

One component. One change. The answer the agent gets before it edits anything.

Or read the public proof first →