← Blog Models & Routing 7 September 2026 10 min read

Benchmarks for your own harness

Public coding benchmarks are under strain: one of the best known was retired by a lab that helped build it, and the newest ones now need a quality team of their own. Those benchmarks measure models. We had a smaller, stranger problem: measuring the app around the model, on your own tasks.

Fig 01  What each kind of benchmark holds stillSchematic
PUBLIC BENCHMARK Which model is better? Fixed task set Fixed agent harness The model (varies) Strain: tasks leak into training data, tests reject correct fixes, scores saturate, agents learn to game the verifier. LUMINAIR EVAL CASE Did this change help on my work? One of your own tasks Its starting snapshot Its acceptance checks The model (varies) The Luminair build (varies) One trial rule (varies) Filled: held still. Dashed: the one thing you change. CAN DRIFT UNSEEN model used per turn · context the harness sent · shared memory and connectors · the verifier files
A drawing of the idea, not of a screen. The bottom strip lists the four things our own audit found could change between two runs without anyone noticing. Everything below is about recording or freezing them.

The short version

  1. In February 2026 OpenAI stopped reporting SWE-bench Verified: saturated, contaminated, and with many tests that reject correct answers. Terminal-Bench now treats its benchmark as software with continuous QA.
  2. Those benchmarks compare models. An app like Luminair also needs to know whether a change to its own harness made your real tasks go better or worse.
  3. On 6 September 2026, any goal that passes its checks started saving itself as a private eval case you can rerun on another model. Our own audit that night said these were task reruns, not benchmarks, and listed why.
  4. Since 7 September a rerun gets its own workspace, records the model used on every turn and a manifest of what was sent, and can freeze rules and tools so exactly one thing differs between two runs.
02Public benchmarks under strain

When the scoreboard wears out.

A benchmark is a fixed set of tasks with an automatic grader. For coding agents, SWE-bench Verified was the standard one for a long time: 500 tasks taken from real open-source repositories, each with tests that decide whether a fix works.

On 23 February 2026, Latent Space interviewed two members of OpenAI's Frontier Evals team about their decision “to publicly abandon SWE-Bench Verified today and endorse SWE-Bench Pro”. The reason, in their words: “the eval is effectively saturated and also highly contaminated.” The write-up adds that problems are sourced from open-source repositories that “many model providers use for training purposes”, and that in a deeper review of problem cases, 49 tests were too narrowly defined, rejecting functionally correct submissions, while 26 looked for features the problem never described.

Terminal-Bench, which tests agents on command-line tasks, took a different path. When it released version 2.0 alongside Harbor, a harness for running agent evaluations, the team wrote that “Substantial manual and LM-assisted verification went into the creation of each task in Terminal-Bench 2.0.” By August 2026, Snorkel's post on Terminal-Bench 4.0 described a whole review pipeline, starting from the problem that “Most benchmarks are static datasets with no active maintenance, causing them to lose value fast.”

One stage of that pipeline is worth quoting in full, because it matters below: “Given the chance, an agent will read the test files, fake the output, or edit the verifier, and it will look like a pass.”

A benchmark is software. It needs maintenance, and it needs to be protected from the thing it measures.
03Why a harness needs its own

Same model, different app.

The harness is everything around the model: the system prompt, the context it injects, the tools it offers, when it retries, how it decides a task is finished. Luminair is a harness around Claude, Codex, Gemini and other engines. When we change it, the models stay the same, but the results on your work can change a lot.

Public leaderboards cannot tell you that. They fix the harness and vary the model. We needed the opposite: fix the task and vary the harness, or fix the harness and vary the model, on tasks that look like what our users actually do. And we did not want anyone's code to leave their Mac to do it.

The natural source of such tasks turned out to be a feature we were building anyway.

04Goals become eval cases

A task with a verdict.

Harness v2 landed on 6 September 2026 (commit 7731ed6e). One of its pieces is the goal. You name an outcome and a command that proves it, for example npm test. After each model turn Luminair runs the command. If it fails, the output goes back to the model as the next prompt, until the command passes or a budget runs out. The model's own “done” is never enough; without a check, the app says in so many words that completion “is a claim, not proof.”

A goal that passes has everything a benchmark task needs: an objective, a grader, and a known starting point, because Luminair's Rewind feature snapshots the folder before every turn. So the same commit made every passed goal save itself as a local eval case. (Today, failed attempts with a check are saved too; a failure is also a result.) The file's opening comment explains the intent: rerun it later “on another model or a newer build of the harness”, and compare. “Compared across runs that is continuous evidence of harness lift, on YOUR tasks, kept local. Nothing here uploads anything.”

Fig 02  What an eval case storesField names from lib/eval-cases.js
The case
objectiveWhat should be completed
gateThe command that must exit 0
checksUp to 12 named acceptance checks
startHashThe Rewind snapshot the work began from
cwdThe project folder
Each result
modelOr mixed: a, b if the run changed model
statuspassed, failed, and so on
tokens · msCost and wall time
harnessThe Luminair version that ran it
manifestFingerprint of what was sent
runManifestOne record per turn in the run
Fields marked “7 Sep” arrived with commit d9ac9038, the day after the audit. Up to 200 cases and 50 results per case are kept, in a JSON file in the app's support folder. The comparison table groups results by model and by environment fingerprint, so runs in different environments are never averaged together.
05The audit said: not yet

A rerun is not a benchmark.

The same night, a second end-to-end audit of Harness v2 went through every new feature, following commands from the composer down to disk. It found twelve issues. Two were about eval cases, and together they explain the difference between rerunning a task and measuring something.

F4

The rerun rewound your folder

To rerun a case, the app restored the project folder to the case's starting snapshot, then opened a new session. If another session or your editor was working in that folder, it was rewound under them. Rewind keeps a recovery snapshot, so nothing was lost for good, but nothing was isolated either.

F12

The result could name the wrong model

The model was captured when the goal started. If the engine switched or fell back mid-run, the result still named the first one. Cases kept no record of memory, tools, dependencies or permissions. And a fresh session still received the project's shared memory, which could include the original answer.

F12 ended with a sentence we took as a rule: “Until then, describe this as a local task rerun, not a rigorous benchmark.” The eval sheet in the app still says so today: “these are task reruns, not controlled benchmarks.”

The memory point is the subtle one. Luminair's Journal remembers things across sessions, which is the point of it. But a task you already solved may have left a note behind. A “fresh” rerun that can read that note is not measuring the model; it is measuring the model plus a hint.

06What a fair rerun needs

Change one thing.

Two commits on the morning of 7 September answered it. The first, f9604d69 at 07:30, gave each rerun its own workspace, reserved it in the main process and started recording the model on every turn. The second, d9ac9038 an hour later, is titled “Add harness acceptance checks, recovery receipts and controlled trial inputs”. Here is how they map onto the findings.

Fig 03  From finding to fixAudit F4, F12 · f9604d69, d9ac9038
BeforeSince 7 September
Shared project folder rewound→The snapshot is exported into a disposable workspace. The project folder is untouched; you review a diff to keep a result.
Goal armed from the interface, after the session started→The run is reserved in the main process and armed on the new session's first event, so a fast model cannot finish first.
Model captured once, at the start→Every turn adds its model. Two or more become mixed: a, b in the result.
No record of what was sent→A manifest per turn: engine, model, each injected block's size and SHA-256, tool names, sandbox policy. No prompt text.
Shared memory in the window→Optional memory and connectors excluded by default. Frozen mode keeps only standing rules, copied and hashed.
No saved snapshot, rerun anyway→Refused: “a controlled rerun is not possible.”
Left column paraphrases the audit note of 7 September 2026 (UTC). Right column is read from the current eval:run handler in desktop/main.js, lib/harness-evidence.js and lib/eval-environment.js.

Frozen inputs

The strictest mode, Freeze rules and SDK tools, builds a profile for the case the first time it runs and pins it: the starting snapshot, a fixed tool list (Read, Write, Edit, Bash, Grep, Glob), no connectors, no skills, no plugins, and only the standing rules that applied at the time. It also pins the runtime: Luminair version, Agent SDK version, platform, processor, Node version, and hashes of the controller file and of our evidence rule. The profile gets a fingerprint.

On every later run the app compares. If the runtime or the sandbox policy differs, it refuses to launch: “The frozen trial runtime changed. This run was not launched.” That sounds strict, and it is. A comparison across two different builds is still possible; it is just not labelled as a frozen one.

Frozen mode also allows one controlled change: a trial rule of up to 8,000 characters, added to that run only. The result is filed under its own fingerprint, labelled with the rule's hash, next to the frozen baseline. That is the experiment you actually want: “does this one instruction help on my task?” For now, frozen trials run only on the built-in Claude models; other engines can run ordinary isolated trials.

Protecting the grader

Remember the Terminal-Bench line about agents editing the verifier. Goals now take up to 12 named acceptance checks, and each can list verifier files to protect. Their SHA-256 hashes are recorded when the goal starts. If a file changes, before or during a check, the goal stops as blocked: “Acceptance verifier changed”, review it and start a new goal to approve it. Symlinks cannot point a verifier outside the project.

The lesson we tookThe hard part of a benchmark is not the score. It is being able to say what was held still. If you cannot list what differed between two runs, you have two anecdotes, not a comparison.
07Find it in the app

Build your own in three steps.

  1. 1In a session, type //goal new. Fill in What should be completed?, then an acceptance check with a Verification command such as npm test. List any test files under Verifier files to protect. Press Start goal.
  2. 2When the goal finishes, it is saved as a case. Type //eval to see every case with its results per model and environment.
  3. 3Pick a model, keep Exclude optional memory and connectors ticked, and press Run isolated trial. For an A/B test of an instruction, tick Freeze rules and SDK tools and paste the rule to test.

Every trial uses your normal account and its usage, and nothing is uploaded.

08What we checked

Checked, and not claimed.

33/33
Goal, roadmap and harness integration tests pass on the current tree
12
Issues in the second Harness v2 audit; two (F4, F12) were about eval cases
6
SDK tools in a frozen trial: Read, Write, Edit, Bash, Grep, Glob
8,000
Maximum characters in a trial rule

What this post does not claim

  • Any harness-lift number. We are describing the instrument, not publishing results from it.
  • That a frozen trial is fully deterministic. Models are not, and engine-owned history and native tools are outside what the manifest measures.
  • That your cases say anything about other people's work. They are built from your tasks, which is both the point and the limit.

Sources

For why the harness matters at all, read Your model didn't get dumber. Your harness might have.

Measure it on your own work

Turn a finished goal into a case, then rerun it on another model in its own workspace.

Download Luminair →