Dying before turn one
108 sessions were abandoned, most before doing any real work. The startup path is the fragile part nobody hardens.
108 sessions — 9.1% of the corpus — were abandoned. Not failed mid-task, not interrupted by the user: abandoned, with gaps of more than 24 hours and no clean ending. And 106 of the 108 died before doing any meaningful work.
Look at where they died: generating the session title. Splicing the agent inbox. Setting up the policy. Starting the first step. The startup path — the unglamorous sequence of setup calls before turn one — is where sessions go to die.
Why it matters
Everyone hardens the runtime and nobody hardens the boot sequence. But from the user's perspective, a session that dies during setup is worse than one that fails mid-task: there's no partial result, no error they can act on, just a session that never started. Multiply by 9% and it's a quiet reliability disaster.
Startup fragility is also invisible to the usual monitoring. Health checks pass — the server is up, the models respond. The failure is in the choreography: five setup steps that each assume the previous one worked, with no verification between them.
What the fix looks like
Treat startup as a pipeline with the same rigor as the runtime. Each setup step gets a success signal, not an assumption. Title generation fails? Fall back to a default and move on — a missing title is cosmetic, not fatal. Inbox splicing fails? Retry with backoff, then degrade to an empty inbox rather than a dead session.
And instrument it. The corpus could tell us exactly where sessions died because the events were logged — but nobody was watching that signal. A startup dashboard that counts 'sessions created vs. sessions reaching turn one' would have surfaced this immediately. The cheapest reliability win in the system is measuring the thing everyone assumes works.
Designing setup that degrades
The fix for each fragile step follows the same pattern: try, fall back, degrade — never die. Title generation fails? Use the first line of the prompt as the title. Inbox splicing fails? Start with an empty inbox and log it. Policy setup fails? Fall back to the most restrictive default policy — deny-by-default is always a safe degraded state — and flag the session for review. The principle: setup steps are conveniences, not prerequisites. The only true prerequisite for turn one is a working model connection, and even that deserves a clear error instead of a silent abandon. Every fallback should be loud in the log and invisible to the user — they get a working session, and the operator gets the signal that setup degraded.