Skip to contentAitium

Blog

The 60-minute invisible stall

A work cell sat stalled for an hour while every health check reported green. Liveness has to be native to the runtime, not a dashboard someone watches.

Somewhere in the directive corpus, a work cell stopped making progress and sat there for 60 minutes. Not crashed — stalled. The process was alive, the server was up, the model endpoint responded to pings. Every health check in the system reported green, for an hour, while nothing happened.

This is the failure mode that dashboards can't catch, because dashboards measure the wrong thing. CPU fine, memory fine, endpoint reachable — all true, all irrelevant. The question isn't 'is the machinery running.' It's 'is the work moving.' Those are different questions, and only one of them was being asked.

Why it matters

A stall is silent money burn with a side of false confidence. The operator glances at the dashboard, sees green, and moves on. Meanwhile the job they needed done isn't done, the deadline slides, and nobody knows until someone asks 'hey, where's the result?' An hour later.

The deeper problem is that external watchers can only see symptoms. A polling script that checks 'is the process alive' will never distinguish 'working hard' from 'stuck.' It sees the same heartbeat either way. The distinction between progress and mere aliveness lives inside the runtime — in the event stream, where turns start and end, where tool calls complete. Only the harness itself can see it.

What the fix looks like

Emit liveness as a first-class signal from inside the session. The runtime already knows when the last real progress happened — the last completed turn, the last finished tool call, the last event that moved the work forward. A heartbeat service watches that stream with an injected clock: if no progress event has landed within the threshold, it emits a stalled signal. When progress resumes, it emits recovered. Exactly once per episode — not a spammy timer, a state change.

Two details matter. First, the signal must be log-only and native — an event in the session's own log, not a row in someone's monitoring database. That way it survives restarts, it's auditable, and any consumer (an alert router, a supervisor agent, a human) can subscribe to the same stream. Second, own signals are excluded: the heartbeat watching itself would be a farce. Progress means the work moved, not that the watcher blinked.

The 60-minute stall would have been a 5-minute stall with this in place. Green dashboards would still say green. But the session log would say stalled, and that's the signal that counts.

← All articles