Skip to contentAitium

Blog

Pass counts don't prove correctness

A green suite can test nothing at all. Prove the control is what makes the case pass — by watching the disabled twin fail.

Here's a test that passes and proves nothing: a test for an admission check that never actually exercises the check. The setup grants everything, the assertion checks the happy path, the suite goes green, and the admission logic could be deleted entirely without a single failure. Pass counts are vanity. What matters is whether the test can fail — and whether it fails for the right reason.

We learned this building governance features, where the whole point is that something gets refused. A test that only checks 'allowed calls go through' tells you nothing about the refusal path. The refusal is the feature. If your tests can't produce a refusal, you haven't tested the feature.

Why it matters

Untestable-by-construction code rots silently. Someone refactors the admission check six months later, the tests stay green, and the governance is gone — nobody knows until an ungoverned call does something expensive. The test suite becomes a comfort object instead of a safety net.

This applies far beyond governance. Any test where the critical behavior is the absence of something — no leak, no double-spend, no unauthorized call — is at risk. Absence is easy to assert and easy to fake. The test passes because nothing happened, and nothing happening is also what a broken test looks like.

What the fix looks like

Disabled-control twins. For every acceptance test, build the twin: the same scenario with the mechanism under test switched off. The real case must pass; the twin must fail — and fail for the right reason, not by accident. If the twin passes, your test doesn't test what you think it tests. If the twin fails for the wrong reason (a crash instead of a clean refusal), your test is asserting on noise.

Concretely: testing a budget rollback? Run the scenario with rollback disabled and watch the budget leak — that's the twin proving the rollback is the control. Testing an admission refusal? Run with the grant present and watch it pass — proving the refusal comes from the missing grant, not from a broken test harness. Each twin is a small experiment with a predicted outcome, and the prediction is the point.

We run these twins on every governance feature, and they've caught real gaps — including three missing rollbacks that the happy-path tests never touched. A test that can't fail is a decoration. Build the twin, watch it fail correctly, then trust the green.

← All articles