The 51-retry burn loop
One session retried a dead transport 51 times. Retry without a circuit breaker isn't recovery — it's spending money to learn nothing.
We scanned 1,186 agent session logs and found a session that retried a failed network call 51 times. Fifty-one. The transport was dead from the first attempt — every retry hit the same wall, and the harness kept paying for each one.
Why it matters
Retries exist because transient failures are real: a 503 here, a timeout there. But a retry policy with no memory is not a recovery strategy. It can't tell the difference between 'the server hiccuped' and 'the server is gone.' So it keeps going, burning tokens and wall-clock time, and the user gets the bill for 51 identical failures.
The corpus backs this up beyond the worst case. Transport failures, timeouts, and 503s accounted for 151 retries across 33 sessions. Every one of those retries was the harness spending the user's money to re-learn something it already knew.
What the fix looks like
Three things, all boring, all necessary. First, backoff with a cap: each retry waits longer than the last, and the count has a hard ceiling — single digits, not fifties. Second, a circuit breaker: after N consecutive failures of the same kind, stop calling and mark the route dead for a cooldown period. Third, classify before retrying: a timeout might be transient, a malformed response never gets better on retry.
The principle is simple. A retry should be a bet with positive expected value — 'this might work now.' After a handful of identical failures, it's not a bet anymore. It's a loop. Break it.
Telling transient from permanent
Not every failure deserves the same treatment, and the retry policy should know the difference. A timeout might clear; a connection refused usually won't; a malformed response never gets better no matter how many times you resend it. Classify first, then decide: retry the maybe-transient, break the certainly-permanent. And make retries idempotent where it matters — a retried write that isn't idempotent can do the thing twice, which turns a burn loop into a correctness bug. The circuit breaker stops the bleeding; classification and idempotency stop the retry from causing a second, different failure.
There's also an organizational lesson in the 51. Nobody noticed. The retries were logged — 151 of them across the corpus — but no alert fired, no dashboard turned red, no budget alarm tripped. A burn loop that nobody watches is just a slow leak with extra steps. Cap the retries, break the circuit, and alert on the break. The operator should hear about the first circuit trip, not discover the fifty-first retry in a log review weeks later.