Stop coding agents from claiming "done" without evidence
Last updated:
A coding agent's "done" is only evidence when the checks you chose ran after its last change and their results are tied to that exact state of the code. DoneLatch, a free MIT tool for Claude Code, Codex, Gemini CLI and Cursor, refuses acceptance until such fresh passing checks exist and a deliberate fault shows the check can actually fail.
Facts on this page were checked against the sources linked inline and the DoneLatch v0.1.2 README on October 3, 2026.
Why "the agent said it passed" is not enough
- In Stack Overflow's 2025 Developer Survey, 66% of developers who answered the AI-frustrations question picked "AI solutions that are almost right, but not quite", the most-cited answer; 45.2% picked "debugging AI-generated code is more time-consuming" (survey). These are survey answers, not a measured failure rate.
- A Claude Code bug report describes the model calling work "verified green" and done without running the project's canonical build-and-test command (anthropics/claude-code#63861, May 30, 2026).
- A Codex report describes background tasks whose final message says a command ran, while the host shows no observable command record or real exit code (openai/codex#34152, July 19, 2026).
Two different gaps show up here. The agent may not have run the check at all. Or it ran a check that cannot fail for the bug you care about.
The three things to require
| Requirement | What it rules out | How DoneLatch does it |
|---|---|---|
| Fresh: the checks ran after the last relevant change | "Tests passed" from before the final edit | Each receipt binds the Git HEAD, watched file names, SHA-256 content hashes, modes and the config hash. Any later change makes it stale. |
| Recorded: results come from the tool, not the agent's summary | A narrative of success with no command behind it | run executes the approved commands itself and writes a signed receipt with their results and output |
| Sensitive: the check can detect a representative fault | A green test that would also pass with the feature broken | faultcheck injects a fault you define into a temporary copy and requires the check to fail with your declared marker and exit code |
The third row is a small negative control, the idea behind mutation testing. DoneLatch's demo shows why it matters. A persistence test that only checks saveSettings() === true still passes when the write is deleted, so DoneLatch refuses "done". A check that reads the saved bytes catches the same fault, and DoneLatch accepts.
Set it up (about a minute, Node.js 24+)
npx --yes github:alidaram99/donelatch#v0.1.2 init # writes receipts.yml (does not run anything)
# Edit receipts.yml: one trusted acceptance command and one meaningful fault.
npx --yes github:alidaram99/donelatch#v0.1.2 trust # HUMAN ONLY: review and type APPROVE <hash>
npx --yes github:alidaram99/donelatch#v0.1.2 run
npx --yes github:alidaram99/donelatch#v0.1.2 faultcheck
npx --yes github:alidaram99/donelatch#v0.1.2 verify-done # exit 0 only for current, sensitive evidenceThen attach it to the agent's stop event, so the agent cannot finish its turn on stale or missing evidence:
| Agent | Hook event | Install |
|---|---|---|
| Claude Code | Stop | claude plugin marketplace add alidaram99/donelatch, then claude plugin install donelatch@donelatch-marketplace |
| Codex | Stop | codex plugin marketplace add alidaram99/donelatch --ref v0.1.2, then install it and trust the hook in /hooks |
| Gemini CLI | AfterAgent | gemini extensions install https://github.com/alidaram99/donelatch --ref v0.1.2 |
| Cursor | stop | manual adapter in the agent install guide |
When evidence is missing or stale, the hook asks the agent for one correction turn. After that it reports the work as UNVERIFIED instead of looping.
Limits
- Humans approve the checks.
trustruns only in an interactive terminal and records the hash of the exact configuration outside the repository. Changingreceipts.ymlrequires a new approval. - The approval does not cover everything. It does not pin the contents of the scripts your checks call,
PATHresolution or the inherited environment. - You choose the fault. DoneLatch does not invent a correctness oracle. A check passes this gate only for the fault you wrote down.
- It is a cooperative guardrail. Hooks can be disabled, skipped or time out, and then they fail open. A same-user agent that deliberately obfuscates or rewrites local files can defeat any hook-based guard, including this one. That class is out of scope. For an adversarial agent, use an OS sandbox, a container or a separate OS user. Enforce acceptance with
verify-donein a CI step the agent cannot edit. - Local signatures are not independent attestation. The same OS user can read the signing key.
FAQ
How do I make Claude Code run the tests before it says it is done?
Add a Stop hook that runs a verifier and blocks the stop when evidence is missing. DoneLatch's Claude Code plugin does this. It requires a passing run and faultcheck for the current files before verify-done accepts.
Is a passing test suite enough evidence?
Not on its own. The tests must have run after the last change, and at least one of them must fail when the behavior you care about is broken. A check that cannot fail is not evidence.
Does this replace code review or mutation testing?
No. It is a small, explicit completion gate for agent workflows. Full mutation testing, independent review and production monitoring still have their place.
Can a determined agent get around it?
Yes, if it runs as your user and sets out to evade the check. Treat the hook as a guardrail for cooperative agents. Use OS-level isolation and a CI acceptance step for agents you do not trust.
Related
- Up: all guides on this site
- Next: stop agents installing packages that do not exist (ExactGround)
- Get DoneLatch (free, MIT) · website