← Blog Engines 27 September 2026 9 min read

Harness engineering, from the inside.

This spring the industry argued about whether the harness, the software around a model, matters as much as the model itself. We build a harness for fifteen engines. Here is what it actually does, what it absorbs, and what it cannot fix.

Fig 01  Many dialects, one vocabularyEvent names from the code
WHAT EACH TOOL PRINTS Codex thread.started turn.started item.completed · agent_message item.started · command_execution turn.completed Antigravity init · conversation_id step_update · text_delta result · status …and thirteen more engine files ONE FILE PER ENGINE engines/<id>.js translates one line at a time into Luminair's words event: (ev) => … ONE VOCABULARY session alive thinking text tool usage done error THE SHARED RUNNER Every engine gets Stop that kills the whole tree Start watchdog and heartbeat Queues and supersede Live mirror on other devices Token spend and meters Crashes as readable sentences Limits and account hops
Raw event names from desktop/engines/codex.js and desktop/engines/antigravity/stream.js; the vocabulary and the runner's duties from desktop/engines/_TEMPLATE.js and runner-cli.js. Layout is schematic.

The short version

  1. A harness is everything around the model: the loop that runs tools, the context it is given, the timeouts, the retries, the rules. The model thinks; the harness decides what it sees and what happens next.
  2. In February and March 2026, OpenAI described building a product with agents and a carefully engineered harness, and Latent Space asked whether the harness is where the value really is.
  3. From inside a harness that runs fifteen engines, our answer is narrower: the harness cannot make a model smarter, but it decides whether each model gets a fair chance, and it is where most day-to-day failures live.
  4. Luminair adds an engine with one file, translates every tool's output into one small vocabulary, and replays 28 recorded traces on every test run to keep them honest.
02Built with Luminair

Every quirk here was found live.

Luminair is built inside Luminair, and its sessions run on whichever engine suits the job: Claude, Codex, Gemini, Antigravity, Kimi and others. Each quirk in section 05 was found the same way: a turn behaved oddly in real work, someone looked at the raw stream, and the fix went into that engine's file with a comment saying what was seen and when. Most of those comments are dated, and the dates below come from them.

Luminair · sessions on different engines
Luminair desktop app in the light theme: a folder sidebar and three chat sessions side by side
Sessions side by side, each on its own model. From here, one engine and another look the same; making them look the same is the harness's job. (Demo workspace.)
03The debate

Big model or big harness?

In February, InfoQ reported on OpenAI's “harness engineering”: in a five-month internal experiment, engineers “built and shipped a beta product containing roughly a million lines of code without any manually written source code”. The engineers' work moved to “designing environments, specifying intent, and providing structured feedback”, with architecture rules enforced mechanically, by linters and structural tests. InfoQ quotes Martin Fowler: “Harness includes context engineering, architectural constraints, and garbage collection.”

In March, Latent Space asked whether harness engineering is real. It framed the tension as “Big Model and Big Harness” and gathered both sides. On one side, Claude Code's Boris Cherny: “all the secret sauce, it's all in the model. And this is the thinnest possible wrapper over the model.” On the other, results where the harness moved scores, and Jerry Liu's line: “The Model Harness is Everything.” The piece also cites a benchmark where one model did better in Claude Code than in a generic agent and another did the reverse.

That last point is the one we recognise. A harness is not neutral. It is tuned, knowingly or not, to the models its authors use most. A multi-model app feels that every day, because the same harness has to serve models that were trained on very different tools.

The harness cannot make a model smarter. It can make sure the model gets a fair chance, the same on every engine.
04One file per engine

Add an engine, restart, done.

Luminair's engines live in one folder, desktop/engines/. Today it holds fifteen: two for Claude (one through the Agent SDK, one through its command-line tool), then Codex, Gemini, Antigravity, Kimi, Meta, Grok, Perplexity, an OpenAI API lane, DeepSeek, GLM, Qwen, Ollama and LM Studio. The loader picks up every file in the folder, skipping templates, shared machinery and launchers. There is no list to edit.

A new engine starts as a copy of _TEMPLATE.js. Its header is a short promise:

desktop/engines/_TEMPLATE.jsheader
// 3. Restart Luminair. That is the whole job. The model picker,
//    the account menu, the "add account" list, badges, colours,
//    per session engine memory, the run, Stop, queues, live
//    mirrors and token meters all pick it up.

Only one field is required: an id. Everything else, from the model list to sign-in to the usage meter, is optional and gets a safe default, so a half-finished engine degrades gracefully instead of crashing the app.

The heart of the contract is translation. Most engines are command-line tools that stream JSON, one object per line, and every vendor names things differently. An engine file declares how to start its tool and one function that reads a single line and answers in Luminair's words: a session id, some text, some thinking, a tool call, a usage count, done or an error. The Codex file, for example, turns thread.started into a session, item.completed with an agent_message into text, and turn.completed into done.

Once a file speaks that vocabulary, the shared runner does the rest: Stop that ends the whole process tree, the start watchdog and heartbeat, queued and superseded turns, the live mirror your phone watches, token spend, and turning a crash into a sentence you can read. That is where most of the harness actually lives, and every engine gets it without writing it.

05What it absorbs

Every engine is different.

The vocabulary is the easy part. The hard part is everything each tool does that no documentation mentions. A selection from the engine files, each with the date its comment gives:

CODEX

Long silence before the first word

On a large resumed thread the model can think for minutes before replying, and a 60 second watchdog killed six such turns in one morning (7 September). The tool's turn.started line is now read as proof of life.

AGY

A cap per conversation

Antigravity refused resumed conversations past roughly 2 MB of history with a quota error, while the same account answered a fresh one that minute. Bracketed on 18 September: 1.96 MB answered, 2.2 MB refused. The runner now rotates to a fresh thread at 1.6 MB.

AGY

Repeated boilerplate

Luminair re-sends its session rules on every turn. On a resumed Antigravity thread those rules were already in the transcript: measured on 18 September, 44% of a 3.6 MB transcript was repeats. Only the parts that change are sent now.

AGY

A keychain that did not exist

Each account runs in its own home folder. Antigravity stores its token through the macOS keychain, and a bare folder has none, so macOS asked about a missing keychain. Each home now gets its own, created once, without touching yours.

KIMI

A login that failed after it worked

After the token lands, Kimi's sign-in still sets up a default model. When Kimi's server answered 500 there (seen 7 September), the tool reported a failed login with a valid token. Luminair retries that step without a browser.

META

Model names checked on the server

Meta's tool rejects an unknown model id outright. So the file only maps ids the real tool has reported, never assumed ones, and the default model sends no id at all.

Even the shape of a tool call differs. When Luminair records what a turn did, it has to find the file or command in whatever field the engine used:

desktop/lib/causal-ledger.jsone line, many dialects
inp.command || inp.CommandLine || inp.file_path || inp.path
  || inp.AbsolutePath || inp.TargetFile || inp.SearchPath
  || inp.SearchDirectory || inp.pattern || inp.Pattern
  || inp.query || inp.Query || inp.description || inp.prompt
  || inp.Prompt || inp.url || inp.Url

None of this makes any model better at coding. All of it decides whether a model gets to finish its turn. A harness that kills a thinking model, or resends the same rules on every turn, or reports a good login as broken, will make a strong model look weak, and the user will blame the model. We wrote about two of those cases in Your model didn't get dumber.

06Harness v2

Make the harness checkable.

On 6 September Luminair shipped what the commit calls Harness v2. Several of its parts answer the harness debate directly, because they make the harness something you can inspect rather than trust:

  • //goal: “A goal the harness finishes, not the model: done only when the gate command passes.” The model does not get to declare victory; a command you choose does.
  • //eval: “Your own benchmark: replay a passed goal on another model and compare pass rate, tokens, time.” The same task, the same harness, a different model.
  • //context, //trail and //doctor: what is in the window, what each prompt caused and cost, and one health pass over the whole harness.

The part that keeps the engines honest is the conformance lab. It replays a recorded, secret-free trace for each engine through that engine's real parser and, for the command-line engines, through the real runner with a fake process. When Harness v2 landed it held 25 traces for 13 engines. Today:

Fig 02  The conformance lab, todayReal data · 55 of 55 pass
Antigravity
Claude CLI
Codex
DeepSeek
Gemini
GLM
Grok
Kimi
LM Studio
Meta
Ollama
OpenAI API
Perplexity
Qwen
normal turnerror turnteardown error after a finished reply
  • Thread id captured, so resume works
  • Tool calls extracted with their names
  • Text survives into the transcript
  • Usage never runs backwards
  • Exactly one terminal event, never two or zero
  • Auth and limit lines are errors, not done
  • Garbage lines never throw
  • Stop kills the child and still ends once
Each block is one recorded trace in desktop/test/fixtures/engine-traces (28 files; Meta's second is its local Glimmer model). The checks are the contract listed at the top of test/core/engine-conformance.test.js. Run on 27 September: 55 tests, 55 passed. The Claude SDK lane is not a command-line tool and has no trace here.

The striped block is a good example of why the lab exists. In long sessions, Antigravity sometimes reports a network error during teardown after the reply has fully streamed and every tool has closed. Treating that as a failure would throw away a finished answer. The parser keeps the completed turn, and a recorded trace makes sure it keeps doing so.

07Find it in the app

Put the harness to the test.

  1. 1In any session, type //goal with a gate, for example //goal --gate "npm test" make the failing tests pass. The goal ends only when the gate passes.
  2. 2Type //eval to list passed goals, and //eval run with a number to replay one on another model and compare.
  3. 3Type //doctor for a health pass over logins, engine tools and connectors, with safe one-tap repairs.

Every // command is listed under Settings › General › Luminair Abilities › View.

08Honest limits

What we can't claim.

What this post does not claim

  • That Luminair's harness makes any model score higher. We have not published benchmark comparisons between Luminair and each vendor's own tool.
  • That every engine behaves identically. The lab proves a shared contract, not equal quality. Some engines have fewer features, and Kimi has no recorded error trace yet.
  • That the quirks above are permanent. Vendors change their tools; a comment dated September describes September.

Sources

One harness, fifteen engines

Pick a model per session, switch mid-thread, and let the harness keep them all honest.

Download Luminair →