Skip to contentAitium

Blog

Budgets that leak when runs fail

A budget reservation that survives a crashed run is a leak. Reserve, commit, roll back — on every failure path, including the ones in your error handling.

Budgets in agent systems work like this: before a run spends anything, it reserves an amount. If the run completes, the reservation commits. If it fails, the reservation rolls back and the budget is freed. Simple — until you enumerate the failure paths. The run can fail. The commit can fail. The audit logging around the commit can fail. Each of those is a path where a reservation can get stuck, and a stuck reservation is a leak: budget that's neither spent nor available, forever.

We found these the hard way. The sweep that reviewed the budget implementation found not one but three missing rollbacks: the lazy stream reservation when stream setup throws, the reviewer reservation when packet audit throws, the sweep reservation when brief audit throws. Three different features, same bug shape — the error path in the error handling had no cleanup.

Why it matters

Leaked reservations are the slow version of the 51-retry burn. Nobody notices one leak. But reservations accumulate across every failed run, and failed runs are the norm in agent systems — roughly 3% of turns in our corpus end in error or abort. Over weeks, the leaked budget quietly strangles the system: agents get refused for budget they should have, operators see 'insufficient budget' for workloads that should fit, and nobody can explain where it went because the leak left no trace.

It's also a trust problem. A budget system that leaks is a budget system the operator stops believing. And once the operator stops believing the budget numbers, they start bypassing them — raising limits, disabling checks — and then you have no budget system at all.

And the stakes aren't theoretical. We've burned 10 billion tokens on a single failed run — ten billion, on work that produced nothing. Across failed projects the total is around 30 billion tokens. That's not a model problem; the models did what they were asked. It's a runtime problem: no budget that could say stop, no admission that could refuse, no rollback that could recover. Every one of those tokens was spendable right up until someone noticed. That's why hardening the agentic layer isn't a nice-to-have — it's the difference between a failed run and a failed run that also burns the budget.

What the fix looks like

Reserve/commit/rollback as a strict protocol, with rollback on every exit path — including throws inside the audit and logging code that wraps the commit. The pattern that caught our three leaks: wrap the reserve in try, and roll back in catch and in finally-adjacent cleanup, then rethrow. The rollback itself must be infallible or at least non-throwing; a rollback that throws is just a new leak with extra steps.

And test the failure paths, not just the happy path. The acceptance test for a reservation isn't 'reserve then commit works.' It's 'reserve then throw at every stage works, and the budget is whole afterward.' Disabled-control twins apply here too: run the same scenario with rollback disabled and watch the budget drain. If the twin doesn't leak, your test isn't testing the rollback.

Boring, mechanical, and the difference between a budget system and a budget rumor.

← All articles