← Blog Building Luminair 27 September 2026 9 min read

A second model reviews the first

In one year, AI code review went from a feature to a product category. We use it on ourselves. Here are two of Luminair's riskiest changes, what a second reader found in each, and the one thing a second model cannot give you.

Fig 01  Three places a second reader sitsFrom commits and review records
DEVICE MIRROR · 7 AUG 2026 Mirror v1 written sessions and folders synced 8-angle adversarial review reviewer model not recorded 18 fixes, one commit including one sync document per device COMPUTER-TO-COMPUTER RELAY · 15 SEP 2026 Written with Claude Fable 5.1 Reviewed by Claude Opus 5 read-only agent, spawned by the implementer 18 findings, 17 fixed and retested commit 2ea95ca · one left open, on the record EVERY FINISHED MASTER CONTROL ITEM · IN THE APP Item marked done by the session that did it Independent feature verifier fresh context, four read-only tools PASS earns “tested” only you can mark it confirmed
The first two lanes are real events from Luminair's history, from commit messages and the review record kept in our Journal vault. The third is a feature in the app, on by default. Shaded boxes are the second reader.

The short version

  1. The model that wrote a change is the worst-placed reader of it. It shares the author's assumptions. A second reader with a fresh context does not.
  2. Luminair's device mirror went through an 8-angle adversarial review that produced 18 fixes, including a redesign that removed a whole class of lost updates.
  3. The computer-to-computer relay, written with Claude Fable 5.1, was reviewed read-only by a Claude Opus 5 agent. It raised 18 findings; 17 were fixed and retested the same day.
  4. The same idea ships in the app as Solace's independent verifier. And our own record is clear that none of this counts as an independent review.
02Review became a product

From feature to category.

In September 2025 OpenAI released GPT-5-Codex, a model tuned for long refactors and for review. As InfoQ reported, on recent commits from popular open-source repositories it “produced review comments that were more accurate and higher-value, reducing noise for developers and highlighting critical issues.”

In March 2026 Anthropic launched Code Review in Claude Code. Cat Wu, Anthropic's head of product, described the reason to TechCrunch: with Claude Code “putting up a bunch of pull requests, how do I make sure that those get reviewed in an efficient manner?” The design is several agents at once, “each agent examining the codebase from a different perspective or dimension”, and a final agent that “aggregates and ranks the findings, removing duplicates”.

By August 2026 CodeRabbit, which reviews AI-generated code, had closed a $143 million Series C. The pattern is clear. When a machine writes most of the code, a human cannot read all of it, so another machine reads it first.

03Why a second reader

The author cannot see its own assumptions.

Ask the model that just wrote a change whether it is correct, and it will mostly re-read its own reasoning. The bug is usually in an assumption it never wrote down: that a timestamp is always present, that a path is what it looks like, that two computers never write at once. A reader who did not make those assumptions has a chance of seeing them.

Luminair is built inside Luminair, and about two in three of its commits carry a Claude co-author line. So for changes that touch other people's computers or other people's data, we add a second pass whose only job is to break the first. Two of those passes are worth walking through, because they both found real problems and both came back with exactly 18 items.

04The mirror: 18 fixes

One writer per document.

The device mirror keeps your sessions and folders in sync across your computers and phone. Version 1 was written in early August 2026. On 7 August it went through what the commit calls an “8-angle adversarial review”, and one commit landed 18 fixes from it. The record does not say which model ran that review, so we will not claim one.

The biggest fix was a change to the contract. Every computer used to write its folder list into one shared cloud document, sync/tree, by reading it, merging its own changes and writing the whole thing back. That is the textbook shape of a lost update.

Fig 02  Why one shared document lost writesSchematic
BEFORE · ONE SHARED DOC Mac A Mac B sync/tree 1 read v1 2 read v1 3 A writes v1 + a 4 B writes v1 + b a is gone, and A's cache says it landed: no retry AFTER · ONE DOC PER DEVICE Mac A tree_A Mac B tree_B Readers list every tree_* document and combine them. One writer per document: nothing to overwrite.
Illustrative sequence; the document names are real (users/{uid}/sync/tree_{deviceId}, version 2). The commit calls the old design “a lost-update trap made permanent by the landed-write suppression cache”.

The fix did not add locks or retries. It removed sharing. Each device now writes only its own document, and readers list all of them and combine them, resolving per-session overrides by timestamp. The commit message puts it in one line: “One writer per doc kills the whole conflict class.”

The other 17 were smaller but just as concrete. A rename made during a network blip used to be lost for good; now a failed forced write retries in 30 seconds. After a restart, a surviving local ledger could bring a deleted session back to life; the first publish now checks the cloud's deletion record first. And the live feed could go “deaf-but-alive” for up to 55 minutes when its token expired mid-stream; it now drops the token on those error frames.

05The relay: 18 findings

“Not shippable as reviewed.”

On 15 September 2026 Luminair gained a computer-to-computer relay: a session runs on the computer that holds its files, and your other computers become live windows into it. That means one Mac accepting prompts from another. The commit that added it carries a Claude Fable 5.1 co-author line.

Later the same day, a Claude Opus 5 agent reviewed the two relay commits read-only. It raised 18 findings. Its verdict, as the review record summarizes it: the worker was “close to shippable with two defects”, and the desktop executor was “not shippable as reviewed”.

Fig 03  The relay review's 18 findingsReal counts · review record, 15 Sep 2026
High Medium Low Info 5 7 4 2 17 fixed in commit 2ea95ca and retested.1 open: a device name taken from the ticket request, which needs a backend deploy.
Severity counts as written in the review record. Bar length is proportional to the count.

All but one were fixed in commit 2ea95ca and retested that afternoon; the review record tracks them as objectives with paired failing and passing tests. Among them: every requester now has a budget per prompt, with a reaper after three hours so a crashed run cannot wedge it; the room allows 120 prompts a minute per account and each executing Mac 40 a minute with 6 running at once; reply frames follow the room's own record of who asked, and a requester binds to the first computer that claims the run, so another device on the account cannot inject text into a pane mid-turn.

06What tests missed

Bugs in the unwritten case.

The relay already had tests when it was reviewed. The review still found bugs, because tests check the cases their author thought of. Two small ones show the pattern well. Both are about input that is missing or not what it looks like.

desktop/main.js · stale prompt checkbefore and after 2ea95ca
// before: a prompt with no timestamp skipped the age check entirely
if (p.ts && Date.now() - p.ts > RELAY_INBOX_MAX_AGE_MS) { … }

// after: no timestamp means infinitely old
function _relayAge(p) { const ts = Number(p && p.ts); return Number.isFinite(ts) && ts > 0 ? Date.now() - ts : Infinity; }
if (_relayAge(p) > RELAY_INBOX_MAX_AGE_MS) { … }   // a doc with no timestamp is stale by definition

The first test anyone writes for a staleness rule uses an old timestamp and a new one. Nobody writes one with no timestamp at all, and that was the case that slipped through: a prompt with the field missing never counted as stale, however long it sat.

The second is folder containment. When another computer asks a Mac to start a new session in a folder, the Mac must refuse unless that folder is one it holds. The old check compared the path as given. The new one requires an absolute path, resolves symlinks and .. with realpathSync on both sides, and only then asks whether one sits inside the other. A path that looks like it is inside your project but points somewhere else no longer passes.

The shape of a good review findingNeither bug needed deep insight. Each needed someone to ask “what if this is missing?” or “what if this is not what it looks like?” about a line the author had stopped seeing.
07The verifier in the app

A fresh reader, on every item.

The same idea is a feature you can use. Solace keeps a ledger of what you asked each session to do. When a session marks an item done, the Independent feature verifier can check it. It is a fresh agent that never saw the original work. It gets the requirement and the change, and its instructions say its job “is NOT code review”; it is to trace every control the requirement introduces from the button to the handler to the state to the code that reads it, and to fail the item if any link is missing.

Three details make it a second reader rather than an echo. It starts with no session history. It can only read: four tools, Read, Grep, Glob and LS, inside the project folder, with no connected tool servers. And a PASS only earns the item the status “tested”. Only you can mark it confirmed.

On 27 September 2026 we fixed what the verifier is shown. It used to get one diff cut at 60,000 characters, which missed new untracked files entirely and, on a busy shared checkout, could cut off the very change it was meant to check. Now new files are included, the files the item's own session changed come first and are labeled, and the budget is shared per file so one huge file cannot hide the rest. Anything not fully shown is named, with an instruction to read it before judging.

08Not independent

A second model is not a second opinion.

Here is the part we want to be careful about. The relay was also assessed against our internal security baseline, and that record lists “Self-review: the implementer assessed its own change” as a finding. The Opus 5 review did not close it. The record says so directly: the review “was run by an agent spawned by the implementer; not independent”, and it is logged with independent: false.

That is the right call. The reviewer was a different model, but the same company's models, briefed by the session that wrote the code, looking at what that session chose to show it. A different model catches different slips. It does not bring different incentives, and it cannot tell you what it was not asked to look at. The relay's assessment still reads incomplete, with independent review named as missing evidence.

So we use a second model the way you would use a sharp colleague on a busy day: it finds real bugs, cheaply and fast, and it is not a substitute for someone accountable signing off.

09Find it in the app

Get a second reader.

  1. 1Open the Solace panel and check that Independent feature verifier is on under Quality gates (it is by default). It runs by itself when an item reaches done; the magnifier button on a done item runs it again.
  2. 2For a quick second opinion on a turn, switch the session to a model from another company. The new model gets a catch-up note of the conversation and can review the last answer.
  3. 3Ask the reviewer to be adversarial and read-only, and to list what it did not check. Then read that list.
10What we checked

Checked, and not claimed.

18
Fixes in the device mirror review commit, a4f62e9
17/18
Relay review findings fixed and retested in 2ea95ca
10/10
Tests pass for the verifier's scope and Solace's plan and proof rules
4
Tools the verifier may use, all read-only

What this post does not claim

  • Which model ran the device mirror review. The record does not say.
  • That either review was independent. The relay record explicitly says it was not.
  • How GPT-5-Codex, Code Review or CodeRabbit would have done on the same code. We did not run them.
  • That the verifier catches every broken feature. It fails when unsure, and it can still be wrong.

Sources

Let a fresh reader check the work

Let the independent verifier check what “done” really means.

Download Luminair →