2026-10-05 · 7 MIN READ
The heartbeat
My merge gate went 52 minutes without merging anything, and nothing said so. By morning the heartbeat my agents built to catch stalls like that had fooled itself twice with one git check, on branches that had no work on them yet.

At four minutes to eleven on the night of October 4, I typed one line into the session that runs my game's build pipeline: get more done, always be working on the game. The session went looking for idle capacity and found something worse. The merge gate had not merged anything in almost an hour. A dozen batches in a row had stopped on the same word, Aborting, and nothing had told either of us.
The game is Spaceframe, and AI agents write most of its code under my direction. One Claude Code session is the coordinator. It writes task briefs and hands them to coding lanes, which are Claude agents plus two copies of a local Qwen model on my own graphics cards. Every finished branch goes to the gate, a script that merges branches in batches and runs the game's probes, automated tests that boot the game and check one thing each, before anything reaches main. The gate has its own earlier post. This one is about the part that is supposed to notice when the gate stops.
The cause was mundane. Just after ten, the disk filled up in the middle of a merge. The gate works in its own disposable copy of the repository, and the failed merge left that copy half-written. Every batch after that tried to check out a branch, hit changed files git would not overwrite, and gave up before running a single probe. A cleanup agent freed the disk less than twenty minutes later, which fixed nothing, because the half-written copy was still there. Worse, the gate marked each aborted branch as tried, so by a quarter to eleven it had run out of branches and gone quiet.
The first aborted batch started at 22:05 and the last at 22:46; nothing merged in between. The gate's copy was reset at 22:57, and that batch merged its first two branches. The 20 are the distinct branches across the 12 aborted batches, all put back in the queue.
Unjamming it took minutes. The coordinator reset the gate's copy to main and put every dropped branch back in the queue, and the first batch after the reset merged two of them. It also made the gate inspect its own copy at the start of every batch and reset it if it was dirty, since that copy holds nothing worth keeping. The harder question was why an hour of failure had raised no alert. The monitor, a script that passes notable lines from the gate and the lanes to the coordinator, did read the gate's log, but it read it for words: PUSHED, RED, CONFLICT. It even listened for ABORTED, which is what the gate prints when it stops a probe early. Git's word was Aborting, and that was not on the list. Nothing was watching for the change that mattered, which was that PUSHED had stopped appearing.
A few minutes later I asked whether System 1 should keep a heartbeat on everything and escalate problems. System 1 is a twelve-billion-parameter Gemma model running on the CPU, which the pipeline already uses to sort failing probe rows into likely causes. The coordinator said yes, with one condition: "as long as System 1 isn't the part that notices silence." A model judges what it is given to read, and a stalled gate gives it nothing new to read.
So the heartbeat has two layers. Every five minutes, plain rules check that each moving part is still moving: each lane's log advances, the model servers answer, the disk has room. For the gate, the log must keep growing while branches wait in its queue, and its recent lines must contain no abort or disk errors. Its very first run, with the jam's lines still in the log, flagged them. Strictly, it still does not watch for PUSHED to stop appearing, so a gate that fails in some new way while writing steadily to its log would get past it. That gap is still open. A problem is reported once when it appears and once when it clears. Only a new problem goes to System 1, which picks a likely cause and a next step from fixed lists (wait, rerun, restart or escalate) and recommends. It acts on nothing. If the heartbeat script crashes, the monitor prints that too.
I widened the brief before it was built. The heartbeat should keep development moving, not just catch breakage, so it also watches for idle capacity: an empty lane queue, a gate with nothing to merge, too few agents working. And it got one action of its own. When a lane's queue is empty and a reserve folder holds briefs, it moves the lowest-numbered brief into the queue. A brief moved too early is cheap to move back, which is why that was the one action I was comfortable automating. A quiet pass takes about two seconds. A Claude agent built and tested it in about eight minutes, with simulated stalls and refills, and it went live before midnight.
Timings from the building agent's test runs on the live pipeline. Passes run every 300 seconds. System 1 shares the CPU with other work, so a slow answer is cut off at the budget and the problem is escalated without a suggestion.
On its first test run it flagged a model server that was down while two lanes were using it, and System 1 called it a dead server and suggested a restart, which was right. A little after two in the morning came a stall it had no rule for. An agent building a toy house for the level editor ran a quick Godot script to inspect its work, and the script never quit. The probe scheduler that lets lanes share the machine treats any running Godot as me playtesting, so it held back every probe that needs the screen or the whole machine, for every lane, behind a script that would never finish.
The coordinator found it by hand at 2:41, after one lane's probe batch had sat for half an hour with an empty log. The hung script had been running for 27 minutes by then, and killing it got the waiting probes moving again within a few minutes. The heartbeat now flags any Godot that has been running a script for more than fifteen minutes, and agents must put a timeout on any Godot they start by hand.
Around five, a different check misfired the other way. The monitor also runs a rules-only lane-health verdict on each working lane, and it flagged one lane as stuck after an hour because the same probe command had failed seven times in a row. Nothing had failed. The probe wrapper exits with a special code that means the batch is still queued and should be run again, and the lane was doing exactly that. The coordinator excluded that command from the repeated-failure rule; on the next check a second rule, one that counts identical repeated calls, called the same lane spinning, and got the same exclusion. A check that cannot tell waiting from stuck teaches its reader to skip it.
Meanwhile the heartbeat had picked up more work than the one action I had allowed. Shortly before midnight I had told the coordinator to stop waiting for my call on design decisions, and two of its decisions gave the heartbeat new jobs. Each probe prints one pass or fail line per check, and the gate lets a branch merge when its only failing lines are on a list of checks already failing on main. Each check on that list names the branch that will fix it. From half past midnight a small list script, run by the heartbeat on every pass, took a check off the list once its fix branch reached main. From a quarter past five it also released briefs that had been waiting on another lane's branch, once that branch reached main. Both jobs decided "reached main" the same way, by asking git whether the branch was an ancestor of main.
At 6:27 the heartbeat printed that it had released a brief because its dependency had merged. A lane had only just created that dependency's branch. A new branch points at main's latest commit, and git counts a commit as its own ancestor, so a branch with no work on it yet passes the test as merged. The coordinator read the line within ten seconds, put the brief back on hold and changed the test. A dependency now counts as merged only when the gate has tested a version of that branch and that tested version is on main.
The coordinator did not think to fix the list script, which used the same test. At 7:21 three water-physics checks went on the list, naming a fix branch that an agent created moments later. About a minute later the heartbeat printed that all three had been taken off because their fix had merged. The fix had not been written. The coordinator read that line six seconds after it arrived, gave the list script the gate-tested rule and put the three checks back. The same test probably explains a check for storm rain that had vanished from the list at 4:21, seconds after its own fix branch was created, and that held three branches back until the coordinator noticed it was gone. No line about that removal ever reached the monitor, though the water removals printed, and I cannot say why, so it stays a likely third case rather than a confirmed one.
Neither confirmed mistake let bad code through. Taking a check off the list makes the gate stricter, not looser, and a brief released early would at worst cost one lane run that ends with the lane reporting it is blocked. Both printed mistakes were caught within seconds, while the unreported failures earlier in the night each ran for close to half an hour or more. I would not lean on that contrast. The coordinator happened to be reading the monitor closely when those lines appeared, and the unreported failures came while it was busy with other work.
This is one night of one pipeline. I have no measured accuracy for System 1's suggestions on heartbeat alerts beyond that first correct call. The time limits on the rules are first guesses, and a heartbeat that hangs instead of crashing would still look quiet: it writes a timestamp on every pass, but nothing reads that timestamp yet.
The part I own is the one automatic action. I allowed exactly one and chose it because a wrong move was cheap. By morning the heartbeat was taking two more, and nobody had held them to that standard, because shortly before midnight I had told the coordinator to stop waiting for me. Both went wrong on the same question, whether everything on a branch is already on main, and for a branch with nothing on it git correctly says yes.
- 01Nothing merges on its own word
The gate this heartbeat watches: why no branch merges until its probes are rerun by something other than the agent that wrote it.
/blog/nothing-merges-on-its-own-word
- 02The harness was the slow part
How the local Qwen lanes the heartbeat keeps fed came to work on their own.
/blog/the-harness-was-the-slow-part
- 03git merge-base --is-ancestor
The check at the center of the twist. It answers whether one commit is reachable from another, which is true for a commit and itself.
https://git-scm.com/docs/git-merge-base
- 04Google SRE book: Monitoring distributed systems
The standard case for alerting on symptoms, such as work that stopped happening, rather than on the specific errors you already thought of.
https://sre.google/sre-book/monitoring-distributed-systems/