The short version
- In February 2026 OpenAI stopped reporting SWE-bench Verified: saturated, contaminated, and with many tests that reject correct answers. Terminal-Bench now treats its benchmark as software with continuous QA.
- Those benchmarks compare models. An app like Luminair also needs to know whether a change to its own harness made your real tasks go better or worse.
- On 6 September 2026, any goal that passes its checks started saving itself as a private eval case you can rerun on another model. Our own audit that night said these were task reruns, not benchmarks, and listed why.
- Since 7 September a rerun gets its own workspace, records the model used on every turn and a manifest of what was sent, and can freeze rules and tools so exactly one thing differs between two runs.
When the scoreboard wears out.
A benchmark is a fixed set of tasks with an automatic grader. For coding agents, SWE-bench Verified was the standard one for a long time: 500 tasks taken from real open-source repositories, each with tests that decide whether a fix works.
On 23 February 2026, Latent Space interviewed two members of OpenAI's Frontier Evals team about their decision “to publicly abandon SWE-Bench Verified today and endorse SWE-Bench Pro”. The reason, in their words: “the eval is effectively saturated and also highly contaminated.” The write-up adds that problems are sourced from open-source repositories that “many model providers use for training purposes”, and that in a deeper review of problem cases, 49 tests were too narrowly defined, rejecting functionally correct submissions, while 26 looked for features the problem never described.
Terminal-Bench, which tests agents on command-line tasks, took a different path. When it released version 2.0 alongside Harbor, a harness for running agent evaluations, the team wrote that “Substantial manual and LM-assisted verification went into the creation of each task in Terminal-Bench 2.0.” By August 2026, Snorkel's post on Terminal-Bench 4.0 described a whole review pipeline, starting from the problem that “Most benchmarks are static datasets with no active maintenance, causing them to lose value fast.”
One stage of that pipeline is worth quoting in full, because it matters below: “Given the chance, an agent will read the test files, fake the output, or edit the verifier, and it will look like a pass.”
Same model, different app.
The harness is everything around the model: the system prompt, the context it injects, the tools it offers, when it retries, how it decides a task is finished. Luminair is a harness around Claude, Codex, Gemini and other engines. When we change it, the models stay the same, but the results on your work can change a lot.
Public leaderboards cannot tell you that. They fix the harness and vary the model. We needed the opposite: fix the task and vary the harness, or fix the harness and vary the model, on tasks that look like what our users actually do. And we did not want anyone's code to leave their Mac to do it.
The natural source of such tasks turned out to be a feature we were building anyway.
A task with a verdict.
Harness v2 landed on 6 September 2026 (commit 7731ed6e). One of its pieces is the goal. You name an outcome and a command that proves it, for example npm test. After each model turn Luminair runs the command. If it fails, the output goes back to the model as the next prompt, until the command passes or a budget runs out. The model's own “done” is never enough; without a check, the app says in so many words that completion “is a claim, not proof.”
A goal that passes has everything a benchmark task needs: an objective, a grader, and a known starting point, because Luminair's Rewind feature snapshots the folder before every turn. So the same commit made every passed goal save itself as a local eval case. (Today, failed attempts with a check are saved too; a failure is also a result.) The file's opening comment explains the intent: rerun it later “on another model or a newer build of the harness”, and compare. “Compared across runs that is continuous evidence of harness lift, on YOUR tasks, kept local. Nothing here uploads anything.”
The case
Each result
mixed: a, b if the run changed modelA rerun is not a benchmark.
The same night, a second end-to-end audit of Harness v2 went through every new feature, following commands from the composer down to disk. It found twelve issues. Two were about eval cases, and together they explain the difference between rerunning a task and measuring something.
The rerun rewound your folder
To rerun a case, the app restored the project folder to the case's starting snapshot, then opened a new session. If another session or your editor was working in that folder, it was rewound under them. Rewind keeps a recovery snapshot, so nothing was lost for good, but nothing was isolated either.
The result could name the wrong model
The model was captured when the goal started. If the engine switched or fell back mid-run, the result still named the first one. Cases kept no record of memory, tools, dependencies or permissions. And a fresh session still received the project's shared memory, which could include the original answer.
F12 ended with a sentence we took as a rule: “Until then, describe this as a local task rerun, not a rigorous benchmark.” The eval sheet in the app still says so today: “these are task reruns, not controlled benchmarks.”
The memory point is the subtle one. Luminair's Journal remembers things across sessions, which is the point of it. But a task you already solved may have left a note behind. A “fresh” rerun that can read that note is not measuring the model; it is measuring the model plus a hint.
Change one thing.
Two commits on the morning of 7 September answered it. The first, f9604d69 at 07:30, gave each rerun its own workspace, reserved it in the main process and started recording the model on every turn. The second, d9ac9038 an hour later, is titled “Add harness acceptance checks, recovery receipts and controlled trial inputs”. Here is how they map onto the findings.
mixed: a, b in the result.Frozen inputs
The strictest mode, Freeze rules and SDK tools, builds a profile for the case the first time it runs and pins it: the starting snapshot, a fixed tool list (Read, Write, Edit, Bash, Grep, Glob), no connectors, no skills, no plugins, and only the standing rules that applied at the time. It also pins the runtime: Luminair version, Agent SDK version, platform, processor, Node version, and hashes of the controller file and of our evidence rule. The profile gets a fingerprint.
On every later run the app compares. If the runtime or the sandbox policy differs, it refuses to launch: “The frozen trial runtime changed. This run was not launched.” That sounds strict, and it is. A comparison across two different builds is still possible; it is just not labelled as a frozen one.
Frozen mode also allows one controlled change: a trial rule of up to 8,000 characters, added to that run only. The result is filed under its own fingerprint, labelled with the rule's hash, next to the frozen baseline. That is the experiment you actually want: “does this one instruction help on my task?” For now, frozen trials run only on the built-in Claude models; other engines can run ordinary isolated trials.
Protecting the grader
Remember the Terminal-Bench line about agents editing the verifier. Goals now take up to 12 named acceptance checks, and each can list verifier files to protect. Their SHA-256 hashes are recorded when the goal starts. If a file changes, before or during a check, the goal stops as blocked: “Acceptance verifier changed”, review it and start a new goal to approve it. Symlinks cannot point a verifier outside the project.
Build your own in three steps.
- 1In a session, type //goal new. Fill in What should be completed?, then an acceptance check with a Verification command such as
npm test. List any test files under Verifier files to protect. Press Start goal. - 2When the goal finishes, it is saved as a case. Type //eval to see every case with its results per model and environment.
- 3Pick a model, keep Exclude optional memory and connectors ticked, and press Run isolated trial. For an A/B test of an instruction, tick Freeze rules and SDK tools and paste the rule to test.
Every trial uses your normal account and its usage, and nothing is uploaded.
Checked, and not claimed.
What this post does not claim
- Any harness-lift number. We are describing the instrument, not publishing results from it.
- That a frozen trial is fully deterministic. Models are not, and engine-owned history and native tools are outside what the manifest measures.
- That your cases say anything about other people's work. They are built from your tasks, which is both the point and the limit.
Sources
- Latent Space · 23 February 2026The End of SWE-Bench Verified
- Terminal-BenchTerminal-Bench 2.0 and Harbor
- Snorkel AI · 28 August 2026Terminal-Bench 4.0: Why Continuous Benchmarks Require Continuous QA
For why the harness matters at all, read Your model didn't get dumber. Your harness might have.
Measure it on your own work
Turn a finished goal into a case, then rerun it on another model in its own workspace.