Why my AI keeps a failure memory
The AI I build software with keeps a memory. When something goes wrong, it writes a short note about what happened and the rule that would have caught it. An index of those notes is read at the start of every session. One line per note, enough to decide whether it’s relevant.
Most notes are specific. But at the top of the index is a short section called “recurring failure family”. It groups lessons that turned out to be the same mistake in different clothes. It carries one instruction: suspect it first.
The family is this: a fact written once, when something was set up, that nothing ever checks again. A number copied into a comment. A threshold set once and never re-measured. A baseline captured when a script was written. Each was true on the day it was written. Nothing re-reads it, so nothing notices when it stops being true.
Here is one instance. It lived in the part of my system whose whole job is to notice.
The health check that passed on nothing
Runway, my harness, runs its jobs on a worker program that updates itself. My AI wrote a health check to run after each release, to confirm the worker had picked up the new build. Both halves of that check had the same flaw, and together they reported a clean deploy that had not happened.
The first half compared against a remembered version. When the AI wrote the script, it pasted in the worker’s version at that moment as the “before” value. The check asked: is the live version different from “before”? If so, print “updated”.
Releases went by. The “before” value went stale. When the AI next ran the check on a deploy, the live version was different from “before”, because it was newer than the stale value, not because it had picked up the new build. The worker was still one release behind. The check printed “updated”.
The second half asked for “the most recent run”. It looked up the release workflow by its display name and took the latest run. The display name resolved to an old, stale workflow, so it returned a run from a commit months old. The script never compared that run to the change being checked. It took no argument at all. It reported that old run as this merge’s green.
Why it failed open
A check shaped like “is this different from what I remember?” doesn’t fail when its memory goes stale. It fails open. The staler the remembered value, the more certainly the comparison passes. A broken check that says “fail” gets fixed. A broken check that says “pass” gets trusted.
The fix is not a fresher constant
Updating the remembered version would only reset the clock. The fix is to ask a question the thing itself can answer, with no memory to keep current:
- The worker’s version string includes the full identifier of the commit it was built from. The check now asserts that it contains the commit being checked. A worker that never updated now fails. Nothing passes by going stale.
- The release run is chosen by that commit, not by “latest”.
- The script takes the commit as an argument and fails loudly without one. A script that can’t name what it’s checking can’t check it.
It also gets a negative control on every run: it refuses a missing argument and a commit that doesn’t exist. And before trusting the run lookup on a new release, the AI proved it matched a known good one.
The tell: a pass condition written as “not equal to”, “most recent”, “changed since”, or any value captured when the check was written. Ask what the check would report if the thing under test had not happened at all. If the answer is “pass”, it isn’t a check.
Why keep a memory at all
Each of these failures looked new at the time. A stale comment. A threshold nobody re-measured. A health check comparing against a remembered version. Read one at a time, each gets its own small fix, and the next one arrives looking new again.
Grouped, they are one pattern with one question: what here was written once and is never checked again? Reading that question at the start of every session is cheaper than learning it again.
Two related notes: a green gate covers only its own paths and an ambiguous outcome proves nothing.