SMERESKI
  1. PROJECTS
  2. SKILLS
  3. RESUME
  4. BLOG
  5. GAMES
  6. CONTACT

POST · 2026-10-05

2026-10-05 · 7 MIN READ

When the test is the bug

Over one weekend, the failures my game's merge gate reported kept turning out to be the tests or the test machine rather than the code. The work became telling the two apart without learning to ignore red.

gamesgodotai-agentsverification
Spaceframe far-view terrain before and after: a mountain dome whose evenly spaced rock bands read as contour rings, then the same view with the bands gone.
AI-assisted text — drafted with Claude from my session logs and code, then reviewed by me before publishingAI-built software — every change described here, including the merge gate and its triage rules, was written by Claude Code or a local model (Qwen) under my coordination and reviewed by Claude CodeAI narration — synthetic voice (Kokoro), not a human recording
AI narration — synthetic voice (Kokoro), not a human recording · listen, or click any word to play from there

At four minutes past seven on Monday morning, a test run on a branch full of new 3D models came back with six failures, and the strangest of them were about size. A berry bush the catalog expected to measure about a meter across measured zero by zero by zero. So did a wrecked hauler shell and a crystal cluster from the ruins set, and the ruins that should have stopped a walking player had no collision shapes at all. Read plainly, the models were broken.

Spaceframe is the survival game I have been building in Godot since July, and most of its code is written by AI agents under my coordination. A coordinating Claude session hands work to other Claude sessions and a local Qwen model, each in its own copy of the game, and nothing they write merges until a gate re-runs their probes, the headless tests that measure the game, and blocks on any failure it can't account for. I wrote about that gate when it was new. Over the weekend of October 3 it ran around the clock, and by Monday much of its red was not coming from the game at all.

It started on Sunday evening. Two load-time probes, one timing how long the game takes to put a player on solid ground and one timing the longest stretch any build step holds the main thread, went red on a branch that had not touched loading. Then on the next branch, and the one after that. Each probe prints the physics rate it saw beside its timings, and that number gave it away: the game is built to step its physics 60 times a second, and these runs were getting about 21.

The machine was crowded. That evening one agent was load-testing a wider pool of probe slots so more agents could test at once, first six and then eight, and the coordinator's note from that night puts the slowdown on eight agents probing together. With that many copies of Godot sharing the processor, a timing probe was measuring the box rather than the branch. Every failing load-time row of the weekend, in the gate's log and in the agents' own probe logs, was taken below 24 updates a second.

Physics rate during the weekend's failing load-time runs (updates per second)
what the game is built for60
where the gate now stops trusting a timing row45
fastest failing run23.8
slowest failing run17.5

Every failing load-time row in the weekend's gate log and the agents' probe logs falls between the two highlighted bars. The same rows report 13.3 to 17.7 seconds from load to standing against an 8-second budget, numbers that describe the machine as much as the game.

Just after midnight the coordinating agent changed the script that reads each run. A load-time row is now carried, recorded but not allowed to block the merge, when that same run measured its physics below 45 updates a second, far above anything the starved runs reached and well short of the 60 the game expects. A slow load measured at full physics rate still fails.

The rule has two holes. On a busy night the gate is blind to load time. And a branch that genuinely chokes the main thread can drag its own physics rate under the line and carry itself through. The backstop is a to-do the agent left itself, to run the load-time probe alone on the main branch at a quiet moment every day, and until that runs I have no load-time measurement from a deliberately quiet machine.

The next red was a different kind of wrong. Many probes pin a measured value so that a change nobody intended shows up as red, and when the game changes on purpose, the pin goes stale. At four in the morning a weather branch brought in a wind field, and the storm-rain probe failed. It exists to catch an old bug where rain fell along the world's fixed down axis instead of toward the planet's surface, and its limit assumed the design's gentle storm lean of about 17 degrees. With real wind, storm rain now leans about 30.

How far storm rain leans from straight down (degrees)
the old probe's limit (the design leaned about 17)about 20
storm rain once the wind field arrivedabout 30
the fixed-axis bug the probe exists to catchabout 61

Converted from the probe's own thresholds and readings, which compare the rain's direction with straight down: 0.94 was the old floor, the storm with wind reads 0.865, and the old bug reads 0.49 at the test spawn. The new floor allows about 37 degrees, and the rain must also match the game's rain formula for the current wind to within about 6 degrees, so the check still sits far from the bug it guards.

Widening the limit alone would have kept the check and lost its point. Instead the probe now checks two things: that rain still falls toward the ground, far from the old bug, and that its direction matches the game's own rain formula for the current wind, so a lean of the wrong size still fails. Ground-texture rows and a far-view conifer had gone stale the same way after deliberate look changes, and each re-pin went in as its own commit naming the change it follows and the value it measured.

Not every long-lived red was the test's fault, though. One failure had sat since Saturday on the gate's list of accepted reds, failures already known on the main branch, with a terse note: hut 14 never repaired. That one was real. A worker carrying wood along the wall opposite the village gate was handed a waypoint straight across the village, the movement code refused every step toward it, and the worker stood at the palisade forever while every carrier behind it stalled and the village's wood stock froze. The fix landed at half past midnight with a new probe row that fails before it and passes after. Around four, a probe that films a raid on the village found another real one: under a long alarm the village sheltered and never recovered. By six a fix had it mobilizing instead.

Then, a little after seven, the zero-sized models. Nothing was wrong with them. Before any probe runs, the test script asks Godot to import new assets, and it stopped that import after a fixed two minutes on the assumption the work was done by then. This branch carried about fifty new models; the import was still writing when it was cut off, and the half-imported models loaded as nothing. With the import allowed to finish, the same size rows passed five minutes later.

The fix had a hole of its own. The first version replaced the fixed cutoff with a wait that ends after 30 quiet seconds, under a ten-minute cap. But Godot's import begins with a scan that writes nothing, so in a copy whose cache was already old, the quiet rule could stop the import before it started. Five minutes later a second commit made the old two minutes a minimum before the quiet check applies. That branch is in the gate's queue as I write this.

An hour after that, the village probe went red on main again, not long after the branch carrying the fix for its real bug had merged, and this time the answer went the other way. Run alone on a clean copy of main, it passed. On that one run the agent filed it as a timing flake under load, and left its entry on the list in case the reading is wrong. The same probe had failed twice for opposite reasons, a bug that froze the village's economy and, apparently, a crowded machine, and neither failure message said which.

Afterward, Claude went back through the gate's log and sorted every new red result from the overnight stretch, from the first batch after eight on Sunday evening to the last one before six on Monday, retries and bisect steps included. There were 23. Eight failed only on timing rows from a starved machine. Ten failed on probes that were stale or wrong: the rain lean, river limits re-pinned after two terrain changes were combined on purpose, a water probe whose staged storm the game's own weather code kept resetting, and two probes counting a tree species that grows on only one special world. Eight plus ten makes 18 of 23 where the code was fine and the probe or the machine was not.

Overnight red results at the merge gate, by what was actually wrong
starved machine, timing rows only8
probe wrong or behind a deliberate change10
something else (docs gap, latency, crash, coverage guard)5

All 23 new red results between the Sunday 8:16 p.m. batch and the Monday 5:38 a.m. batch, with retries and bisect steps counted separately. A result lands in the middle bar if any of its rows needed a probe fix or a re-pin; one of those runs was also starved. A row counts only if the accepted list did not hold it at that moment, so real bugs already on the list do not appear here.

That leaves five, and they are part of why the gate still blocks on red. Two runs carried a real gap, three validation rules missing from the mod documentation; the others were a networking row slightly over its latency budget, an unexplained crash and a guard asking to be told about a new far-view hook. The real game bugs of the weekend are not in the count, because the gate only counts reds it has not seen before: the stuck worker was already on the accepted list, and the alarm stalemate came from a filmed raid, not a merge.

The count also bore on a question the agents had been working on all weekend: how to make the gate faster. Two attractive ideas had been measured before anyone built them. Running many small probes inside one Godot process would have saved six or seven seconds per gate batch, and sharing one planet boot across the probes that need a world would have saved a median of 28, against a bar of 60 written down before the numbers came in. The agent that measured the second idea wrote that the bigger wins were elsewhere: in bisects, the re-runs the gate does to find which branch in a batch broke, chasing reds that were not real, and in the crowded machine.

Seconds a speedup would save per gate batch, measured before building
many small probes in one Godot process6 to 7
one planet boot shared by world probes (median)28.3
the bar the shared boot had to clear60

The single-process estimate came to about 2 to 3 percent of probe time; the shared planet boot to 6.9 percent of gate time across 73 batches. The 60-second bar was written into the shared-boot spec before its numbers came in.

Both are parked. What shipped were two plain changes: probes that never render now run without graphics in their own pool, and an agent's batch runs the compile check and the cheap logic probes first, so a failure there skips the slow world-booting probes behind it. On one quiet ten-probe batch that took the time from 515 seconds to 318. The gate itself skips nothing, because it needs every row to sort its reds, so that number is an agent's gain rather than the gate's.

So the lesson I take is triage, not distrust. A red row has three usual suspects, the code, the machine or the pin, and the evidence is often in or near the run: the physics rate it measured, whether it fails alone on main, the commit that last moved the pin. Sometimes it is the bookkeeping. Overnight, the script that retires accepted reds once their fix merges quietly dropped the rain and water rows early, because a brand-new fix branch with no commits looks already merged to the obvious git check. It now counts only a branch tip the gate has actually tested.

The limits are real. This is one weekend on one game, and the sorting is Claude's reading of logs and commit messages, checked against the code wherever a commit claimed a probe fix. Counting retries lets one stubborn branch count several times, and one of the ten stale-probe runs was also starved. Until the daily quiet run reports, I cannot say whether the carried load-time rows are hiding a real slowdown.

The models in that Monday-morning run were never broken. On a branch that size, two minutes was not enough.

REFERENCES5 LINKS
  • 01
    Nothing merges on its own word

    How the merge gate in this post works, and why no agent's own report of a pass counts.

    /blog/nothing-merges-on-its-own-word

  • 02
    The harness was the slow part

    The same lesson from the other direction: measure what a worker is being given before judging the worker.

    /blog/the-harness-was-the-slow-part

  • 03
    The contract came first

    Where the probes-with-thresholds approach, and the pins this post re-checks, come from.

    /blog/the-contract-came-first

  • 04
    Godot command line tutorial

    The headless and import flags the test script uses, including the import step that was cut off at two minutes.

    https://docs.godotengine.org/en/stable/tutorials/editor/command_line_tutorial.html

  • 05
    Godot ProjectSettings: physics ticks per second

    The fixed physics rate, 60 by default, that the load-time probes report beside every timing.

    https://docs.godotengine.org/en/stable/classes/class_projectsettings.html#class-projectsettings-property-physics-common-physics-ticks-per-second