Catching stalls without a dashboard
Exactly-once-per-episode stalled and recovered signals, emitted natively by the runtime. No human watching required.
The 60-minute invisible stall taught us that external monitoring can't see the difference between 'working hard' and 'stuck.' The fix isn't a better dashboard — it's moving the detection inside the runtime, where the event stream lives. A heartbeat service watches the session's own progress events with an injected clock, and when the work goes quiet past a threshold, it says so: natively, as an event, in the session log.
Why it matters
Dashboards fail at stall detection for a structural reason: they observe symptoms from outside. Process alive? Yes. Endpoint responding? Yes. CPU doing something? Probably. None of those answer 'is the work moving.' Only the runtime knows what progress looks like — completed turns, finished tool calls, events that advance the job — because only the runtime defines it.
There's also a scaling argument. A human watching a dashboard works for five sessions. It doesn't work for five hundred, and it definitely doesn't work at 3 AM. Stall detection has to be automatic or it doesn't exist. And 'automatic' means the system itself notices, not a cron job that greps logs.
What the fix looks like
A session-owned heartbeat with four properties. First, it consumes the session event stream directly — no separate monitoring pipeline to keep in sync. Second, it emits exactly once per episode: one stalled signal when progress stops, one recovered signal when it resumes. Not a timer spamming 'still stalled' every minute — state changes, not noise. Third, the signals are log-only session events, so they survive restarts and any consumer can subscribe: an alert router, a supervisor, a human reading the log later. Fourth, the heartbeat excludes its own signals from the progress calculation. A watcher that counts its own blinking as progress is a perpetual-motion machine.
The threshold is the honest engineering tradeoff in the design. Too short and every slow tool call pages someone; too long and you're back to the 60-minute problem. Make it configurable per session, default it sanely, and record what you chose. The point isn't the perfect threshold — it's that the question 'is this stuck?' finally has a native answer instead of a human guess.
Stalls will still happen. That's fine. What's not fine is not knowing.
One caution: heartbeats detect quiet, not wrongness. An agent confidently doing the wrong thing — calling the wrong tool in a tight loop — is making progress by every event-based measure. Stall detection catches 'stopped,' not 'lost.' Catching 'lost' is the claim-audit problem, not the liveness problem. Know which failure each mechanism covers.