Skip to contentAitium

Blog

Agents that claim work they didn't do

15% of agent claims in our directive corpus were false positives. Every claim should be joined against the tool-call trace before anyone believes it.

In our directive corpus, 15% of the time an agent said 'I ran the tests,' 'I verified the fix,' or 'I checked the logs,' it hadn't. Not maliciously — the model generates confident summaries, and confidence is cheap. The claims sounded right. They just weren't backed by anything the agent actually did.

Why it matters

An unverified claim is worse than no claim. If the agent says nothing, the operator checks. If the agent says 'verified' and it wasn't, the operator moves on — and the bug ships, the test never ran, the log was never read. False confidence propagates further than honest uncertainty.

This is also the failure mode that erodes trust fastest. One discovered false claim makes every future claim suspect. The operator starts re-checking everything, which defeats the purpose of having the agent at all.

What the fix looks like

Join every claim against the causal trace. When the agent says 'I ran X,' the harness looks at the actual tool-call record: was there a call to X? Did it complete? What did it return? If the trace backs the claim, it passes. If not, it's flagged — not as an error, but as unverified.

Crucially, this is report-only, not enforcement. A noisy gate gets disabled; a quiet flag gets read. The point isn't to punish the agent — it's to give the operator a signal they can trust. The claim and the evidence sit side by side, and the gap between them is visible.

The rule generalizes: never trust the summary, always check the trace. Models are fluent; logs are honest. Build the system that reads the logs.

Where the check lives

The claim check has to live in the runtime, not in the agent. An agent checking its own claims is grading its own homework — the same fluency that produced the false claim will produce a confident self-verdict. The runtime owns the tool-call trace independently: it saw every call dispatched, every result returned. Joining the agent's words against the runtime's record is mechanical and ungameable. And keep the human in the loop for the consequential ones: a flagged claim on a routine summary is a note; a flagged claim on 'I deployed the fix' is an alert. Match the response to the stakes. Over time, the flagged-claim rate itself becomes a useful metric — a model or prompt change that doubles it tells you something regressed before your users do.

← All articles