A green gate covers only its own paths
A regression gate draws a line around the code it drives. Green means nothing changed inside that line. It says nothing about the rest. The dangerous change is the one whose only visible effect lands outside the line. The gate passes, everyone stops looking, and the sentence the gate seems to support is false.
What happened
My AI was rewriting part of Runway, my harness, while the old version kept running. To keep the old version safe, it had a gate: a set of captured outputs from real runs of the old scripts. Any change that altered what the old scripts printed would turn it red.
One change added a single line to the top of the old planning script. The line pulled in a shared helper script. That looked harmless. But pulling in the helper also installed its exit handler. From then on, if the old planner crashed, it would print a failure marker it had never printed before. And the part of the system that sorts failures into kinds matched on that marker alone. Every crash of the old planner would have been filed under the wrong kind.
The gate stayed green. Its capture of the planner was a successful run that exited early. The exit handler only mattered on a crash. No fixture crashed, so no fixture could see it. Every capture matched, and the old version’s behaviour had still changed.
How to look for it
For any change, ask one question: on which path does its effect show up, and does any test drive that path? Not “did the gate stay green”. It will.
The paths fixtures tend to miss:
- Crash paths. Exit handlers, error handlers, anything that runs on a non-zero exit. Fixtures capture runs that worked.
- Flag-gated branches. The “not a dry run” branch of a release script, the one that actually merges and deploys, often has no fixture at all.
- Retry and repair loops. The second time round is not the first.
- Anything that needs a credential, a daemon or a network the test environment doesn’t have.
State the claim in the gate’s own terms
When a gate is your evidence, say exactly what it proves. “The old scripts print the same bytes on the successful paths we captured” is true and useful. “The old version is unchanged, and the gate proves it” overclaims. That overclaim is what lets a change like the exit handler through, because once the gate is green nobody asks the next question.
The same thing, in time
Repeating a run has the same blind spot. A few clean runs cover only the failure modes those runs varied.
One test checked a trace for a short tool name with a loose pattern. It matched the name anywhere, including inside the random name of a temporary folder the test suite made. When the random name happened to contain those letters, the check failed for no real reason.
The AI had asked for several clean runs in a row, to rule out two AI agents running tests at the same time. The first batch came back clean by luck: none of those runs drew a colliding folder name. A second, separate set of runs caught it.
That bar was set against contention. Contention is nearly deterministic: remove the other process and it’s gone. A random collision is not. Repeating a run a few times says almost nothing about a failure that fires rarely.
The real evidence was forcing it. The AI created a temporary folder with the colliding name on purpose and watched the check fail. That is proof. The repeated runs are the weaker support, not the other way round.
So calibrate the bar to the failure. For contention, remove the competitor and run once cleanly. For a random trigger, force the collision, or run far more times. And when someone offers a green run as evidence, ask what it varied, not how many times it ran.
Chasing that small coincidence turned up a bigger problem. A group of checks meant to prove a tool was never called were anchored on the exact way a call is printed at the top level of a trace. A real call is printed one level deeper. Those checks had passed since the day they were written, and they had asserted nothing.