SMERESKI
  1. PROJECTS
  2. SKILLS
  3. RESUME
  4. BLOG
  5. GAMES
  6. CONTACT

POST · 2026-09-29

2026-09-29 · 5 MIN READ

The harness was the slow part

A local model looked too slow and too lost to work on its own, so I measured what it was being fed. On every turn it read about 30,000 tokens of instructions written for a different model, and every task started in the wrong folder.

aiprocessverificationgames
Spaceframe's galactic map: a web of thin white lanes linking glowing cyan and orange star systems on a black background, with a faint purple nebula and ring labels reading Stone Age, Bow, WW2, Modern and Space
AI-assisted text — drafted with Claude, published autonomously per my standing instruction not to gate a post on my own pre-reviewAI-built software — every change described here was written by a local model (Qwen) or by Claude Code, and reviewed by Claude Code
AI narration — synthetic voice (Kokoro), not a human recording · listen, or click any word to play from there

At three minutes past eight on the morning of September 29, a language model running on one of the two graphics cards under my desk stopped searching my home folder and wrote a report instead of code. It had spent three dozen steps looking for the game it was supposed to change, and its conclusion was blunt: the folder it had been started in was not the project described in its instructions. I read that as one more sign the model was lost.

The game is Spaceframe, the survival game I have been building since July, and most of its code is written by Claude, my coding agent, working against probes I insist on. Claude is also what I pay for by the token. Qwen 3.8 is an open model small enough that one copy fits on a 16 GB card, so two of them run side by side on hardware I already own and cost nothing per use. The day before, I had told Claude to trust Qwen with more of the work, and then that Qwen should do most of it, with Claude reviewing and fixing what came back.

The first job I gave it was the largest: read the whole codebase and report bugs. It went through all 907 source files in about an hour and a half, in parallel on both cards, and came back with hundreds of findings. Claude then checked the 137 it had rated most serious against the actual code, and tallied what held up.

Qwen's whole-codebase bug hunt, checked by Claude
critical or high findings in game code137
confirmed real by Claude4

Claude re-read each finding against the code. Three of the four real bugs were fixed; the fourth looked like deliberate design and was left alone. The 2 to 3 percent hit rate covers open-ended bug hunting only.

Almost none of it held up. The quotes it cited were real lines of code, because a filter threw out any that were not, but the claims attached to them were mostly wrong. It flagged checks the code already made, bugs that comments showed had been fixed months ago, and arithmetic that was simply correct. Nothing checked its work until Claude re-read every line, so every wrong finding cost the paid model time to disprove.

So the next jobs had the opposite shape. Claude pulled the exact lines to change and described the goal, Qwen wrote the change, and something mechanical judged it: a compile check, a probe with a threshold, or a rendered frame that Claude looked at. The game's rain puddles were a good test, because a probe measures how round their outlines are and the wider soft edge I had asked for made them too round. Each round, Qwen got the measured result of its previous attempt and tried again.

Puddle roundness by round (lower is less round; the limit is 0.172)
before0.185
round 10.181
round 20.127
round 30.111

Round 1 only raised one term and barely moved it. Round 2 added a larger lobe shape and passed, but its bigger puddles read too dark from the road. Round 3 kept the shape but was still slightly dark; a final reflection tweak fixed that.

The second round cleared the roundness limit but made the puddles too dark, and the third kept the new shape without quite fixing the brightness. One more pass got it: Qwen raised the puddles' sky reflection, and Claude nudged the same number the rest of the way. The crops went the same way: the game's crops were six-sided cones, and Qwen wrote a 245-line generator for real plants in under seven minutes. It used two Godot functions that do not exist, which Claude fixed and turned into rules. Its second round, after Claude described what the render looked like, is what shipped.

What I actually wanted was an agent: Qwen reading the code, deciding what to change and committing it on its own, the way Claude works. That part kept failing. Of the first six tasks it ran that way, two produced small correct edits and never committed them, and four ran past an hour without changing a file. A watchdog Claude had written kept flagging them as stuck. The obvious conclusion was that a model this size, on a card this size, is too slow to work alone.

Claude measured instead. It sent one trivial request, reply with the word OK, through the setup Qwen had been using and then through a stripped-down one, and compared how much text the model had to read first.

Tokens read before answering OK
Claude Code with my full setup30,040
bare mode2,271

The same one-line request sent to the same model. The full setup also offered 69 tools; bare mode offers three.

The setup Qwen had been using was Claude Code itself, loaded with my whole personal configuration: the global rules I wrote for Claude, descriptions of hundreds of installed skills, an index of my memory notes, and 69 tools. Every step of every Qwen session re-read all of that before it reached the task, which made each step slow. It also meant Qwen was following instructions written for a different model, such as rules for how Claude should plan and report. Claude Code has a bare mode that skips all of it, so Qwen now gets three tools and its own short playbook instead.

To test that fairly, Qwen was given a task whose correct answer we already had: the puddle change from the patch rounds, run again from the original code. The old setup had spent about two hours and twenty minutes on the same brief and left its edits uncommitted.

One task with a known answer, in minutes
full setup (edits left uncommitted)about 140
bare mode (exact answer, committed)13.6

Twenty steps and nineteen tool calls in bare mode. The full-setup time comes from the earlier attempt at the same brief, which also started in the wrong folder and shared its card with other jobs for part of the run, so it overstates what bare mode alone saved.

In bare mode it went straight to the right lines, made the edits, reviewed its own diff, noticed it had missed one and fixed it, then compiled and committed. Its change matched the known-good answer exactly. Its steps averaged about 41 seconds, against roughly 90 in the old setup.

The first real task after that is the one from the start of this post, and it wandered anyway. The program that hands Qwen its tasks, which Claude wrote, started every session in my home folder rather than the game's. Qwen had spent its earlier stuck hours searching my disk for a project it had been told was right in front of it. One missing argument fixed that. The 8:03 report had said exactly this, and I had read it as confusion.

With both problems fixed, it started doing real work on its own. It merged every star on the galaxy map into a single draw call, and Claude fixed one mistake: the stars came out pale because vertex colors are read as linear light unless told otherwise. It also found why a fog test in one of my probes had never worked. The sky system recalculates fog every frame, so the test's tripled fog was overwritten before a single frame was drawn. That is a real root cause, and it is the kind of investigation Qwen had never finished before.

Then I had Claude review everything Qwen had written, eight changes in all. It found nothing critical or high-severity. It found two medium problems, and only one was Qwen's: a refactor that recalculated the same route search once for every planet instead of once in total. The other was a probe check Claude had written itself, which compared a function with another function built on top of it, so it could never fail. Both are fixed.

The part that should compound is the training. Every mistake Qwen makes becomes a numbered rule in a playbook it reads before each task, with a note naming the failure that caused it. There are twenty now. Rule 16 came from a backlog audit that ran out of time with an empty report, because Qwen had planned to write all its findings at the end. Split into smaller pieces with a rule to write each result as soon as it is decided, the next part finished quickly and correctly flagged long-finished tasks that were still listed as open.

The limits matter. These were small, well-specified tasks in one project, over two days, and a sample this small is a set of observations rather than a benchmark. The poor bug-hunting number describes open-ended review, not implementation. Every visual judgment was made by Claude looking at rendered frames, not by Qwen, and my own playtest is still the last word on how the game feels. Qwen is slower per step than a hosted model and still needs a reviewer, but the reviewer now mostly reads diffs instead of doing the work.

I should have measured what the model was being given before I judged it. It was reading someone else's instructions, in someone else's folder.

REFERENCES5 LINKS