Skip to contentAitium

Blog

Context death: finding the limit by hitting it

41 sessions died on context limits the harness never saw coming. The limit should be refused before dispatch, not discovered via a provider error.

41 sessions — 3.5% of the corpus — died because they exceeded a model's context window. In every case, the harness discovered the limit the same way: it sent the request, the provider returned a 400 or 500 error, and the turn died. The harness had all the information needed to know the request was too large. It just never checked.

Why it matters

A context death is the most wasteful failure in the system. The work that built up that enormous context — dozens of turns of tool calls, reasoning, retrieved documents — is all stranded. The session can't continue, can't compact after the fact, can't hand off cleanly. It just stops, mid-thought, and everything it was carrying is gone.

Worse, the failure arrives disguised as something else. Provider 400s and 500s trigger the retry logic, so the harness may burn several more calls re-sending a request that can never succeed. The 51-retry problem and the context problem compound each other.

What the fix looks like

Admission before dispatch. The harness meters every request — prompt tokens plus the response budget — against the model's real window before sending anything. If it doesn't fit, the call is refused with a clear reason, and the session compacts or sheds context first.

This is a reserve-then-commit pattern: check the budget, reserve it, then spend it. The check has to live in the harness, not in the provider's error response. A provider error is a discovery mechanism for the provider's bugs, not your capacity planning.

One session in the corpus hit provider 400s on three consecutive turns — turns 3, 4, and 5 — before anyone noticed the pattern. Three turns of user-facing failure for a limit the harness could have enforced in microseconds. That's the gap.

What to shed first

When the budget check fails, something has to give — and 'something' should be decided by policy, not panic. The usual order: first, drop cached tool results that can be re-fetched; second, compact old turns into summaries, oldest first; third, shed retrieved documents, keeping only what the current step references. What's never shed silently: the task definition and the latest user instructions. A session that compacted away what it was supposed to be doing is technically alive and functionally lost. And every compaction should itself be a logged event — future-you debugging a weird session will want to know what the agent forgot and when.

← All articles