The short version
- A harness is everything around the model: the loop that runs tools, the context it is given, the timeouts, the retries, the rules. The model thinks; the harness decides what it sees and what happens next.
- In February and March 2026, OpenAI described building a product with agents and a carefully engineered harness, and Latent Space asked whether the harness is where the value really is.
- From inside a harness that runs fifteen engines, our answer is narrower: the harness cannot make a model smarter, but it decides whether each model gets a fair chance, and it is where most day-to-day failures live.
- Luminair adds an engine with one file, translates every tool's output into one small vocabulary, and replays 28 recorded traces on every test run to keep them honest.
Every quirk here was found live.
Luminair is built inside Luminair, and its sessions run on whichever engine suits the job: Claude, Codex, Gemini, Antigravity, Kimi and others. Each quirk in section 05 was found the same way: a turn behaved oddly in real work, someone looked at the raw stream, and the fix went into that engine's file with a comment saying what was seen and when. Most of those comments are dated, and the dates below come from them.
Big model or big harness?
In February, InfoQ reported on OpenAI's “harness engineering”: in a five-month internal experiment, engineers “built and shipped a beta product containing roughly a million lines of code without any manually written source code”. The engineers' work moved to “designing environments, specifying intent, and providing structured feedback”, with architecture rules enforced mechanically, by linters and structural tests. InfoQ quotes Martin Fowler: “Harness includes context engineering, architectural constraints, and garbage collection.”
In March, Latent Space asked whether harness engineering is real. It framed the tension as “Big Model and Big Harness” and gathered both sides. On one side, Claude Code's Boris Cherny: “all the secret sauce, it's all in the model. And this is the thinnest possible wrapper over the model.” On the other, results where the harness moved scores, and Jerry Liu's line: “The Model Harness is Everything.” The piece also cites a benchmark where one model did better in Claude Code than in a generic agent and another did the reverse.
That last point is the one we recognise. A harness is not neutral. It is tuned, knowingly or not, to the models its authors use most. A multi-model app feels that every day, because the same harness has to serve models that were trained on very different tools.
Add an engine, restart, done.
Luminair's engines live in one folder, desktop/engines/. Today it holds fifteen: two for Claude (one through the Agent SDK, one through its command-line tool), then Codex, Gemini, Antigravity, Kimi, Meta, Grok, Perplexity, an OpenAI API lane, DeepSeek, GLM, Qwen, Ollama and LM Studio. The loader picks up every file in the folder, skipping templates, shared machinery and launchers. There is no list to edit.
A new engine starts as a copy of _TEMPLATE.js. Its header is a short promise:
// 3. Restart Luminair. That is the whole job. The model picker, // the account menu, the "add account" list, badges, colours, // per session engine memory, the run, Stop, queues, live // mirrors and token meters all pick it up.
Only one field is required: an id. Everything else, from the model list to sign-in to the usage meter, is optional and gets a safe default, so a half-finished engine degrades gracefully instead of crashing the app.
The heart of the contract is translation. Most engines are command-line tools that stream JSON, one object per line, and every vendor names things differently. An engine file declares how to start its tool and one function that reads a single line and answers in Luminair's words: a session id, some text, some thinking, a tool call, a usage count, done or an error. The Codex file, for example, turns thread.started into a session, item.completed with an agent_message into text, and turn.completed into done.
Once a file speaks that vocabulary, the shared runner does the rest: Stop that ends the whole process tree, the start watchdog and heartbeat, queued and superseded turns, the live mirror your phone watches, token spend, and turning a crash into a sentence you can read. That is where most of the harness actually lives, and every engine gets it without writing it.
Every engine is different.
The vocabulary is the easy part. The hard part is everything each tool does that no documentation mentions. A selection from the engine files, each with the date its comment gives:
Long silence before the first word
On a large resumed thread the model can think for minutes before replying, and a 60 second watchdog killed six such turns in one morning (7 September). The tool's turn.started line is now read as proof of life.
A cap per conversation
Antigravity refused resumed conversations past roughly 2 MB of history with a quota error, while the same account answered a fresh one that minute. Bracketed on 18 September: 1.96 MB answered, 2.2 MB refused. The runner now rotates to a fresh thread at 1.6 MB.
Repeated boilerplate
Luminair re-sends its session rules on every turn. On a resumed Antigravity thread those rules were already in the transcript: measured on 18 September, 44% of a 3.6 MB transcript was repeats. Only the parts that change are sent now.
A keychain that did not exist
Each account runs in its own home folder. Antigravity stores its token through the macOS keychain, and a bare folder has none, so macOS asked about a missing keychain. Each home now gets its own, created once, without touching yours.
A login that failed after it worked
After the token lands, Kimi's sign-in still sets up a default model. When Kimi's server answered 500 there (seen 7 September), the tool reported a failed login with a valid token. Luminair retries that step without a browser.
Model names checked on the server
Meta's tool rejects an unknown model id outright. So the file only maps ids the real tool has reported, never assumed ones, and the default model sends no id at all.
Even the shape of a tool call differs. When Luminair records what a turn did, it has to find the file or command in whatever field the engine used:
inp.command || inp.CommandLine || inp.file_path || inp.path || inp.AbsolutePath || inp.TargetFile || inp.SearchPath || inp.SearchDirectory || inp.pattern || inp.Pattern || inp.query || inp.Query || inp.description || inp.prompt || inp.Prompt || inp.url || inp.Url
None of this makes any model better at coding. All of it decides whether a model gets to finish its turn. A harness that kills a thinking model, or resends the same rules on every turn, or reports a good login as broken, will make a strong model look weak, and the user will blame the model. We wrote about two of those cases in Your model didn't get dumber.
Make the harness checkable.
On 6 September Luminair shipped what the commit calls Harness v2. Several of its parts answer the harness debate directly, because they make the harness something you can inspect rather than trust:
- //goal: “A goal the harness finishes, not the model: done only when the gate command passes.” The model does not get to declare victory; a command you choose does.
- //eval: “Your own benchmark: replay a passed goal on another model and compare pass rate, tokens, time.” The same task, the same harness, a different model.
- //context, //trail and //doctor: what is in the window, what each prompt caused and cost, and one health pass over the whole harness.
The part that keeps the engines honest is the conformance lab. It replays a recorded, secret-free trace for each engine through that engine's real parser and, for the command-line engines, through the real runner with a fake process. When Harness v2 landed it held 25 traces for 13 engines. Today:
- Thread id captured, so resume works
- Tool calls extracted with their names
- Text survives into the transcript
- Usage never runs backwards
- Exactly one terminal event, never two or zero
- Auth and limit lines are errors, not done
- Garbage lines never throw
- Stop kills the child and still ends once
The striped block is a good example of why the lab exists. In long sessions, Antigravity sometimes reports a network error during teardown after the reply has fully streamed and every tool has closed. Treating that as a failure would throw away a finished answer. The parser keeps the completed turn, and a recorded trace makes sure it keeps doing so.
Put the harness to the test.
- 1In any session, type //goal with a gate, for example
//goal --gate "npm test" make the failing tests pass. The goal ends only when the gate passes. - 2Type //eval to list passed goals, and //eval run with a number to replay one on another model and compare.
- 3Type //doctor for a health pass over logins, engine tools and connectors, with safe one-tap repairs.
Every // command is listed under Settings › General › Luminair Abilities › View.
What we can't claim.
What this post does not claim
- That Luminair's harness makes any model score higher. We have not published benchmark comparisons between Luminair and each vendor's own tool.
- That every engine behaves identically. The lab proves a shared contract, not equal quality. Some engines have fewer features, and Kimi has no recorded error trace yet.
- That the quirks above are permanent. Vendors change their tools; a comment dated September describes September.
Sources
- InfoQ · 21 February 2026OpenAI introduces harness engineering: Codex agents power large-scale software development
- Latent Space · 5 March 2026[AINews] Is harness engineering real?
One harness, fifteen engines
Pick a model per session, switch mid-thread, and let the harness keep them all honest.