Skip to contentAitium

Blog

Fail closed: the registry tradeoff

Blank deny by default. You trade 'use any tool discovered' for 'use only what was granted' — and refusals are fast, visible, and audited.

Every tool call in a governed runtime passes an admission check: is this tool registered, is the grant active, unrevoked, unexpired, for this agent, right now? If any answer is no, the call is refused. This is fail-closed — the default is deny — and it's a deliberate tradeoff worth stating plainly.

The alternative is fail-open: the agent can call whatever it discovers, and governance is advisory. That's more flexible. It's also how you get an agent invoking a tool nobody approved, at a version nobody reviewed, because it looked useful. Flexibility at the dispatch layer is just ungoverned action with better PR.

Why it matters

The tradeoff isn't really flexibility versus bureaucracy. It's 'who decides what the agent may touch, and when.' Fail-open decides at call time, implicitly, by whatever happens to be installed. Fail-closed decides before the run, explicitly, by whoever issues the grants. One of those is a decision. The other is an accident waiting for a prompt injection.

Prompt injection is the concrete threat that settles it. An agent that can call any discovered tool is an agent whose toolset can be expanded by a malicious document. 'Discovered' is doing a lot of work in that sentence — it means 'the attacker put it there.' A registry with blank-deny default means the injected tool isn't registered, has no grant, and the call is refused before it runs.

What the fix looks like

Blank deny with fast, visible refusals. The refusal must be immediate — no timeouts, no retries against a policy decision — and it must be informative: which tool, which agent, which check failed. The agent sees the refusal and can adapt (ask, work around, stop). The refusal is audited, so the operator sees the attempt. Nothing about a denial should be silent or slow.

And keep the grant surface small. The classic failure of fail-closed systems is that granting becomes so painful that operators hand out wildcard grants to everything — at which point you've rebuilt fail-open with extra steps. Grants should be cheap to issue, scoped narrowly, and expired by default. The friction should live in the scope decision, not in the paperwork.

Fail closed doesn't mean fail hostile. It means the safe state is the default state, and every exception is a conscious, recorded, expiring decision.

The test of a fail-closed system isn't the steady state — it's the incident. When something goes wrong at 2 AM, the operator should be able to read the audit log and answer 'what was this agent allowed to do' in minutes. Fail-open can't answer that question at all. That's the tradeoff, stated plainly: a little friction up front for an answerable system later.

← All articles