Field notes · 2026-09-27

An ambiguous outcome proves nothing

A test that checks an outcome, like a status code, an exit code, an error message or an empty list, has a hidden weakness. If more than one thing in the system can produce that outcome, the test can’t tell them apart. It passes when the thing you care about works. It also passes when that thing is missing and something else produced the same result. That test is green for the wrong reason, and you won’t notice, because it’s green.

What happened

My AI was building a web API. One endpoint was owner-only. It must never accept a caller who hasn’t signed in. The spec asked for proof: start the service with sign-in configured, send an unauthenticated request to that endpoint, and check the answer is 401 (not signed in), not 200 (welcome).

The point was to observe a guarantee that had only ever been assumed.

We have a house rule for checks like this. A check that something is refused needs a companion check that shows the test is looking in the right place. So I asked the AI drafting the change to add one. Send the same request to a misspelled path. That should come back 404 (no such page). If the misspelled path gives 404 and the real path gives 401, the 401 must have come from the endpoint’s own sign-in rule.

The misspelled path also returned 401.

A piece of shared code sat in front of every path under the API’s prefix. Whenever it couldn’t work out who the caller was, it answered 401 and stopped. It did that whether or not the endpoint required sign-in, and whether or not the endpoint existed at all. A request to a path that didn’t exist never got far enough to be told so.

Two separate gates, one answer. The 401 proved a gate was closed. It couldn’t prove which. And the whole reason for the test was to prove the endpoint’s own sign-in rule was attached.

The fix: check the structure, not the outcome

The drafter replaced the misspelled-path trick with a direct check against the service’s route table. It asserts that a route with this exact path and method exists, and that it carries the owner-only sign-in rule. That one check proves the route is real and the rule is attached. It doesn’t reason around the confusion. It skips it.

The status-code check stayed, as a second, weaker observation. Its ambiguity is now written in the test’s own comment instead of being papered over.

The general move: when an outcome has more than one possible source, check the thing that is supposed to produce it. The route table, the registered handler, the rule object. Each of those has one source.

How to spot it before it bites

Ask: what else in this system could produce this exact result? Not “does my code produce it”. The green test already answered that.

Places to be suspicious:

The companion must be independent

The house rule was right: a check that something is absent or refused needs a companion that shows the search really happened. But this incident added a condition. The companion must not share the confusion. Here, the companion was defeated by the same shared code that made the first check ambiguous.

A companion that reads the same kind of signal as the first check isn’t independent evidence. Pick a different kind of observation, not a different value of the same one.

This is a close relative of a green gate covers only its own paths. There, the test never reached the path. Here, it reached the path and couldn’t tell who answered.