I Audited My Agent's Honesty. It Was 5%.

A coverage audit found that 95% of an AI agent's proposed actions had no handler to execute them. The fix wasn't a smarter prompt. It was generating the agent's vocabulary from the code that actually runs.

Five percent.

That’s the number I got the first time I ran a coverage audit against a queue of LLM-proposed actions: enumerate what the agent proposed, cross-reference it against the handler registry, count the mismatches. Roughly 5 of every 93 proposals in the queue would have actually executed if a human approved everything sitting in it.

The queue looked healthy. Every proposal had evidence, a title, a reason. Nothing in the UI said otherwise. The agent was proposing forty things and could do four, and the interface had no way to tell you that.

The failure, concretely: give a model a handler called attach_item_to_shelf and it will, with total confidence, propose reorder_shelf, merge_shelves, create_shelf_page. Plausible verbs, sibling actions, nothing the code can run. A human reviewer approves one. The dispatcher looks for a handler, finds none, logs “skipped,” and the UI shows green. Approved. Nothing happened. That’s the worst version of this bug: not a crash, a silent no-op dressed as success.

A smarter prompt doesn’t fix this reliably. What held was generating the agent’s vocabulary straight from the handler registry, at prompt-build time, so the prompt is structurally incapable of naming an action the code can’t run. Delete a handler, it vanishes from the prompt in the same commit. Add one, the model learns about it with zero prompt edits. There’s no second list to forget to update, because there’s no second list.

Then validate twice, and count the two failure modes separately. Unknown action type is one bug (the model invented a sibling verb). Right verb, wrong param shape is a different bug: I found one action type emitted with eighteen distinct param shapes across a few weeks of runs, title where the schema wanted name, an item name where it wanted an id. Vocabulary validation catches none of the second kind. You need both gates, and you need to know which one is failing, because the fix for each is different: one is a prompt problem, one is a schema problem.

The audit itself is the part worth stealing even if nothing else is. One CLI command, three inputs: the queue, the registry, a status filter. Output is a coverage percentage per topic, a table of unregistered action types with counts, a table of param mismatches with the actual violation string attached. Wire --fail-under 0.8 into CI and a PR that adds proposal-generating logic without the matching handler fails the build before anyone sees a green checkmark that means nothing.

I built this into approval-loop, a small library for the layer every agent framework skips: the gap between “the LLM called a tool” and “the system actually changed.” If you’re running anything that proposes actions for a human to approve, run the audit before you trust the queue. The number will probably surprise you.