Skip to contentAitium

Blog

Reconstructing what an agent actually did

Every admission, refusal, reservation, and claim should join into one reconstructable log. If you can't replay the run, you don't understand it.

When something goes wrong in an agent system — a bad action, a blown budget, a claim that turned out false — the first question is always the same: what actually happened? And in most systems, the honest answer is: we can't fully say. The tool calls are in one log, the policy decisions in another, the budget events somewhere else, and joining them after the fact is archaeology.

The corpus made this concrete. Sessions with hundreds of turns and tens of thousands of events, and the story of any given run has to be assembled from fragments. The events exist — that's the good news — but 'the events exist somewhere' is not the same as 'the run is reconstructable.'

Why it matters

Reconstruction is the foundation everything else stands on. Debugging needs it: you can't fix what you can't replay. Accountability needs it: 'who allowed this call' must have an answer with a timestamp. Claim verification needs it: joining a claim against the tool-call trace only works if the trace is complete and ordered. Even the heartbeat's stall detection is a small act of reconstruction — 'what was the last progress, and when.'

There's also a regulatory shadow here. As agents touch more consequential systems, 'show me exactly what the agent did and why each step was allowed' stops being a nice-to-have. A system that can't produce that record will be excluded from the rooms where the interesting work happens.

What the fix looks like

One event log per session, append-only, with every governance-relevant event in it: admissions granted and refused, grants issued and revoked, budget reserved/committed/rolled back, claims made and their audit verdicts, stalls detected and recovered. Each event carries the session id, a sequence number, and a timestamp. The log is the system of record — not a sidecar, not a best-effort export.

Two properties make it work. First, log-only signals: governance events are recorded as events, never as hidden state transitions. If it wasn't logged, it didn't happen — and if it happened, it's in the log. Second, the log is written by the runtime, not by the agent. The agent's claims go in as claims (marked as such); the runtime's observations go in as observations. The join between them — claim versus trace — is then mechanical, not interpretive.

Reconstructability is a testable property: hand someone a session id and a log, and they should be able to narrate the run — every action, every permission, every dollar — without asking the agent what it remembers. Agents forget. Logs don't.

← All articles