Stop coding agents from claiming "done" without evidence

Last updated:

A coding agent's "done" is only evidence when the checks you chose ran after its last change and their results are tied to that exact state of the code. DoneLatch, a free MIT tool for Claude Code, Codex, Gemini CLI and Cursor, refuses acceptance until such fresh passing checks exist and a deliberate fault shows the check can actually fail.

Facts on this page were checked against the sources linked inline and the DoneLatch v0.1.2 README on October 3, 2026.

Why "the agent said it passed" is not enough

Two different gaps show up here. The agent may not have run the check at all. Or it ran a check that cannot fail for the bug you care about.

The three things to require

RequirementWhat it rules outHow DoneLatch does it
Fresh: the checks ran after the last relevant change"Tests passed" from before the final editEach receipt binds the Git HEAD, watched file names, SHA-256 content hashes, modes and the config hash. Any later change makes it stale.
Recorded: results come from the tool, not the agent's summaryA narrative of success with no command behind itrun executes the approved commands itself and writes a signed receipt with their results and output
Sensitive: the check can detect a representative faultA green test that would also pass with the feature brokenfaultcheck injects a fault you define into a temporary copy and requires the check to fail with your declared marker and exit code

The third row is a small negative control, the idea behind mutation testing. DoneLatch's demo shows why it matters. A persistence test that only checks saveSettings() === true still passes when the write is deleted, so DoneLatch refuses "done". A check that reads the saved bytes catches the same fault, and DoneLatch accepts.

Set it up (about a minute, Node.js 24+)

sh
npx --yes github:alidaram99/donelatch#v0.1.2 init        # writes receipts.yml (does not run anything)
# Edit receipts.yml: one trusted acceptance command and one meaningful fault.
npx --yes github:alidaram99/donelatch#v0.1.2 trust       # HUMAN ONLY: review and type APPROVE <hash>
npx --yes github:alidaram99/donelatch#v0.1.2 run
npx --yes github:alidaram99/donelatch#v0.1.2 faultcheck
npx --yes github:alidaram99/donelatch#v0.1.2 verify-done # exit 0 only for current, sensitive evidence

Then attach it to the agent's stop event, so the agent cannot finish its turn on stale or missing evidence:

AgentHook eventInstall
Claude CodeStopclaude plugin marketplace add alidaram99/donelatch, then claude plugin install donelatch@donelatch-marketplace
CodexStopcodex plugin marketplace add alidaram99/donelatch --ref v0.1.2, then install it and trust the hook in /hooks
Gemini CLIAfterAgentgemini extensions install https://github.com/alidaram99/donelatch --ref v0.1.2
Cursorstopmanual adapter in the agent install guide

When evidence is missing or stale, the hook asks the agent for one correction turn. After that it reports the work as UNVERIFIED instead of looping.

Limits

FAQ

How do I make Claude Code run the tests before it says it is done?

Add a Stop hook that runs a verifier and blocks the stop when evidence is missing. DoneLatch's Claude Code plugin does this. It requires a passing run and faultcheck for the current files before verify-done accepts.

Is a passing test suite enough evidence?

Not on its own. The tests must have run after the last change, and at least one of them must fail when the behavior you care about is broken. A check that cannot fail is not evidence.

Does this replace code review or mutation testing?

No. It is a small, explicit completion gate for agent workflows. Full mutation testing, independent review and production monitoring still have their place.

Can a determined agent get around it?

Yes, if it runs as your user and sets out to evade the check. Treat the hook as a guardrail for cooperative agents. Use OS-level isolation and a CI acceptance step for agents you do not trust.