The short version
- The model that wrote a change is the worst-placed reader of it. It shares the author's assumptions. A second reader with a fresh context does not.
- Luminair's device mirror went through an 8-angle adversarial review that produced 18 fixes, including a redesign that removed a whole class of lost updates.
- The computer-to-computer relay, written with Claude Fable 5.1, was reviewed read-only by a Claude Opus 5 agent. It raised 18 findings; 17 were fixed and retested the same day.
- The same idea ships in the app as Solace's independent verifier. And our own record is clear that none of this counts as an independent review.
From feature to category.
In September 2025 OpenAI released GPT-5-Codex, a model tuned for long refactors and for review. As InfoQ reported, on recent commits from popular open-source repositories it “produced review comments that were more accurate and higher-value, reducing noise for developers and highlighting critical issues.”
In March 2026 Anthropic launched Code Review in Claude Code. Cat Wu, Anthropic's head of product, described the reason to TechCrunch: with Claude Code “putting up a bunch of pull requests, how do I make sure that those get reviewed in an efficient manner?” The design is several agents at once, “each agent examining the codebase from a different perspective or dimension”, and a final agent that “aggregates and ranks the findings, removing duplicates”.
By August 2026 CodeRabbit, which reviews AI-generated code, had closed a $143 million Series C. The pattern is clear. When a machine writes most of the code, a human cannot read all of it, so another machine reads it first.
The author cannot see its own assumptions.
Ask the model that just wrote a change whether it is correct, and it will mostly re-read its own reasoning. The bug is usually in an assumption it never wrote down: that a timestamp is always present, that a path is what it looks like, that two computers never write at once. A reader who did not make those assumptions has a chance of seeing them.
Luminair is built inside Luminair, and about two in three of its commits carry a Claude co-author line. So for changes that touch other people's computers or other people's data, we add a second pass whose only job is to break the first. Two of those passes are worth walking through, because they both found real problems and both came back with exactly 18 items.
One writer per document.
The device mirror keeps your sessions and folders in sync across your computers and phone. Version 1 was written in early August 2026. On 7 August it went through what the commit calls an “8-angle adversarial review”, and one commit landed 18 fixes from it. The record does not say which model ran that review, so we will not claim one.
The biggest fix was a change to the contract. Every computer used to write its folder list into one shared cloud document, sync/tree, by reading it, merging its own changes and writing the whole thing back. That is the textbook shape of a lost update.
users/{uid}/sync/tree_{deviceId}, version 2). The commit calls the old design “a lost-update trap made permanent by the landed-write suppression cache”.The fix did not add locks or retries. It removed sharing. Each device now writes only its own document, and readers list all of them and combine them, resolving per-session overrides by timestamp. The commit message puts it in one line: “One writer per doc kills the whole conflict class.”
The other 17 were smaller but just as concrete. A rename made during a network blip used to be lost for good; now a failed forced write retries in 30 seconds. After a restart, a surviving local ledger could bring a deleted session back to life; the first publish now checks the cloud's deletion record first. And the live feed could go “deaf-but-alive” for up to 55 minutes when its token expired mid-stream; it now drops the token on those error frames.
“Not shippable as reviewed.”
On 15 September 2026 Luminair gained a computer-to-computer relay: a session runs on the computer that holds its files, and your other computers become live windows into it. That means one Mac accepting prompts from another. The commit that added it carries a Claude Fable 5.1 co-author line.
Later the same day, a Claude Opus 5 agent reviewed the two relay commits read-only. It raised 18 findings. Its verdict, as the review record summarizes it: the worker was “close to shippable with two defects”, and the desktop executor was “not shippable as reviewed”.
All but one were fixed in commit 2ea95ca and retested that afternoon; the review record tracks them as objectives with paired failing and passing tests. Among them: every requester now has a budget per prompt, with a reaper after three hours so a crashed run cannot wedge it; the room allows 120 prompts a minute per account and each executing Mac 40 a minute with 6 running at once; reply frames follow the room's own record of who asked, and a requester binds to the first computer that claims the run, so another device on the account cannot inject text into a pane mid-turn.
Bugs in the unwritten case.
The relay already had tests when it was reviewed. The review still found bugs, because tests check the cases their author thought of. Two small ones show the pattern well. Both are about input that is missing or not what it looks like.
// before: a prompt with no timestamp skipped the age check entirely if (p.ts && Date.now() - p.ts > RELAY_INBOX_MAX_AGE_MS) { … } // after: no timestamp means infinitely old function _relayAge(p) { const ts = Number(p && p.ts); return Number.isFinite(ts) && ts > 0 ? Date.now() - ts : Infinity; } if (_relayAge(p) > RELAY_INBOX_MAX_AGE_MS) { … } // a doc with no timestamp is stale by definition
The first test anyone writes for a staleness rule uses an old timestamp and a new one. Nobody writes one with no timestamp at all, and that was the case that slipped through: a prompt with the field missing never counted as stale, however long it sat.
The second is folder containment. When another computer asks a Mac to start a new session in a folder, the Mac must refuse unless that folder is one it holds. The old check compared the path as given. The new one requires an absolute path, resolves symlinks and .. with realpathSync on both sides, and only then asks whether one sits inside the other. A path that looks like it is inside your project but points somewhere else no longer passes.
A fresh reader, on every item.
The same idea is a feature you can use. Solace keeps a ledger of what you asked each session to do. When a session marks an item done, the Independent feature verifier can check it. It is a fresh agent that never saw the original work. It gets the requirement and the change, and its instructions say its job “is NOT code review”; it is to trace every control the requirement introduces from the button to the handler to the state to the code that reads it, and to fail the item if any link is missing.
Three details make it a second reader rather than an echo. It starts with no session history. It can only read: four tools, Read, Grep, Glob and LS, inside the project folder, with no connected tool servers. And a PASS only earns the item the status “tested”. Only you can mark it confirmed.
On 27 September 2026 we fixed what the verifier is shown. It used to get one diff cut at 60,000 characters, which missed new untracked files entirely and, on a busy shared checkout, could cut off the very change it was meant to check. Now new files are included, the files the item's own session changed come first and are labeled, and the budget is shared per file so one huge file cannot hide the rest. Anything not fully shown is named, with an instruction to read it before judging.
A second model is not a second opinion.
Here is the part we want to be careful about. The relay was also assessed against our internal security baseline, and that record lists “Self-review: the implementer assessed its own change” as a finding. The Opus 5 review did not close it. The record says so directly: the review “was run by an agent spawned by the implementer; not independent”, and it is logged with independent: false.
That is the right call. The reviewer was a different model, but the same company's models, briefed by the session that wrote the code, looking at what that session chose to show it. A different model catches different slips. It does not bring different incentives, and it cannot tell you what it was not asked to look at. The relay's assessment still reads incomplete, with independent review named as missing evidence.
So we use a second model the way you would use a sharp colleague on a busy day: it finds real bugs, cheaply and fast, and it is not a substitute for someone accountable signing off.
Get a second reader.
- 1Open the Solace panel and check that Independent feature verifier is on under Quality gates (it is by default). It runs by itself when an item reaches done; the magnifier button on a done item runs it again.
- 2For a quick second opinion on a turn, switch the session to a model from another company. The new model gets a catch-up note of the conversation and can review the last answer.
- 3Ask the reviewer to be adversarial and read-only, and to list what it did not check. Then read that list.
Checked, and not claimed.
What this post does not claim
- Which model ran the device mirror review. The record does not say.
- That either review was independent. The relay record explicitly says it was not.
- How GPT-5-Codex, Code Review or CodeRabbit would have done on the same code. We did not run them.
- That the verifier catches every broken feature. It fails when unsure, and it can still be wrong.
Sources
- InfoQ · 22 September 2025OpenAI Releases GPT-5-Codex Optimized for Complex Code Refactoring and Code Reviews
- TechCrunch · 9 March 2026Anthropic launches code review tool to check flood of AI-generated code
- SiliconANGLE · 12 August 2026CodeRabbit bags $143M to help companies get a grip on the explosion of AI-generated code
Let a fresh reader check the work
Let the independent verifier check what “done” really means.