← Blog Models & Routing 23 September 2026 9 min read

Shipping a new model without shipping an app.

New models now arrive every few weeks. Luminair reads a signed live catalog and can pin a newer vendor CLI, so most new models reach the picker with no app update. Claude Opus 5.5 showed why both halves matter, and where the trick stops working.

Fig 01  Two halves of a new modelSchematic of the code path
BACKEND config/ modelCatalog engines.<id>.models[ ] rollout · kill · plan engines.<id>.cli tarball · sha512 · bin SIGNED · ED25519 HALF 1 · THE ROW Verify signature baked public key Rollout gate bucket 0 to 99 Model in the picker no app build HALF 2 · THE RUNNER Download tarball registry.npmjs.org only sha512 before unpack files only, no links cli/<engine>/<version> preferred over bundle BOTH ARRIVE A model that answers Exception: the built-in Claude lane runs the CLI bundled with the Agent SDK. No pin reaches it.
Drawn from desktop/lib/model-catalog.js, desktop/lib/cli-updater.js and the catalog code in desktop/main.js. The dashed box is the part that still needs an app release, and it is the part Opus 5.5 hit.

The short version

  1. A new model needs two things on your machine: a row in the model picker, and a runner (the vendor's command-line tool) new enough that the vendor will accept it.
  2. Luminair fetches a model catalog every hour. It is signed with Ed25519, rolled out to 10, 50 and then 100 percent of installs, and has a kill switch per row.
  3. The same catalog can pin a newer vendor CLI. The app downloads it from the npm registry, checks its sha512 hash before unpacking, and prefers it over the bundled copy.
  4. Opus 5.5 needed Claude Code 2.1.280 or newer. Our built-in Claude lane runs the CLI bundled in the app, so that one took a release: Luminair 1.0.183, the day after launch.
02The pace

A new model every few weeks.

On 22 September 2026, Anthropic introduced Claude Opus 5.5. Its announcement says it “performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.” Input and output are $4 and $20 per million tokens. TechCrunch noted that the launch “comes just two months after the release of Opus 5 on July 24.”

Our own git history shows the same rhythm from the app side. Opus 5 was added to Luminair's Claude engine on 29 July, Fable 5.1 on 1 September, and Opus 5.5 on 23 September. On the OpenAI side, GPT-6 Astra landed in early September. That is four new top models in the picker inside eight weeks, and each one is a moment where users ask, reasonably, “why can't I pick it yet?”

If every new model waits for an app build, then signing, notarizing, uploading and waiting for every Mac to install the update, you are always a few days behind. So on 7 September we built a way around it. The file that holds the logic opens with the goal in plain words: “a new model appears the moment it is released, with no app update”.

03A row and a runner

The picker is the easy half.

Think of a new model like a new TV channel. You need it in the channel guide, and you need a receiver that can decode it. Put it in the guide without the decoder and you get a blank screen with a nice label.

Luminair runs most engines by driving the vendor's own command-line tool: Codex for OpenAI, Gemini CLI for Google, Claude Code for Anthropic. Vendors can gate a new model on the version of that tool. GPT-6 Astra is where that bit us. The comment on its row in desktop/engines/codex.js records it: the Codex CLI we bundled, 0.147, “got ‘requires a newer version of Codex’ from OpenAI”. It ran only with 0.153 or newer, verified on 7 September with 0.153.4.

That is why the catalog has two kinds of entry. A model row carries only what the picker needs: an id, a label, a short name, a colour and a rank. A CLI pin names an exact build of the vendor's tool, per platform. Both are optional, and a new model often needs both.

Only advertise what runsWhen a vendor's error says this login cannot run this model, or that the tool is too old, Luminair remembers it per account and model for 24 hours. A model blocked on every login of an engine is hidden from the picker. A “needs a newer CLI” error also triggers an immediate catalog refresh, in case a newer build is already pinned.
04The signed catalog

Anything that isn't signed is ignored.

A remote list that can add models, and download programs, is a tempting target. So the catalog is signed. The backend holds an Ed25519 private key; the app ships with the public half baked into its code. The served body is a small envelope: a version, a key id, the catalog as a string, and a signature over that exact string.

desktop/lib/model-catalog.jsverifySignedCatalog, abridged
// Signing the exact string (not a re-serialized object) keeps verification
// independent of key order and number formatting.
const key = crypto.createPublicKey(publicKeyPem);
ok = crypto.verify(null, Buffer.from(body.payload, 'utf8'), key,
                   Buffer.from(body.sig, 'base64'));
if (!ok) return null;

Signing the string rather than the parsed object is a small choice that avoids a class of bugs: two JSON encoders can disagree about key order or how to write a number, and then an honest catalog fails to verify.

The app fetches the catalog three seconds after launch and then every hour. The last good body is cached on disk and re-verified on every load, so the first engine list at boot is already right, even offline, and a tampered cache is treated exactly like a tampered download: ignored. Because the public key lives in the code, rotating it means a new app build. That is the price of trusting nothing else.

05Rollout and the kill switch

Ten percent, then fifty, then everyone.

A signed row is trusted, but not yet proven. Each row can carry a rollout percentage. Every install hashes its device id together with the model's key into a number from 0 to 99, and sees the row only if that number is under the rollout. Because the bucket never changes for a given install and model, raising the number only ever adds people. Nobody sees a model appear, vanish and reappear between refreshes.

Fig 02  One model, 100 install bucketsSchematic
Stage 110%
Stage 250%
Stage 3100%
Kill0%
sees the modeldoes not yetkilled, for everyone
Each cell is a bucket from 0 to 99, and which ones are lit is illustrative. Buckets under the rollout number see the row, so every bucket lit at 10 percent stays lit at 50. The stages and the kill rule come from the backend's rollout step; the order here is simplified.

The backend's hourly watcher moves a healthy row one stage at a time, 10, 50, 100, waiting at least an hour per stage. It reads anonymous telemetry: per model, a hashed install id, the outcome, and no prompt or account. If at least 20 turns ran in the last two hours and more than half failed, it sets kill and raises an alert that the row was rolled back. A killed row disappears for everyone, admins included. CLI pins go through exactly the same gate.

A model the watcher discovers for a command-line engine has to pass a canary on the admin's Mac, which has real logins, before it is promoted to 10 percent: a chat reply and a tool call, each asked to echo a random token like OWCANARY-…. Passing is a string match, not a judgment. Only chat plus tool on at least one login counts as a pass.

06Pinned CLIs

Downloading a program, carefully.

A CLI pin says: for this engine, on this platform, use this npm tarball, with this sha512 integrity hash, and the program lives at this path inside it. desktop/lib/cli-updater.js then does the install in a way that fails closed.

01

Only the registry

A pin whose tarball URL is not on registry.npmjs.org is refused, both when the pin is read and again before download.

02

Hash first

The bytes are checked against the pinned sha512 before anything is unpacked. A mismatch throws and nothing is written outside a scratch folder.

03

Files only

Unpacking keeps regular files and folders only: no symlinks, no absolute paths, no paths climbing out with “..”.

04

Swap in whole

Everything happens in a .partial folder, then one rename moves it to cli/<engine>/<version>. A crash mid-unpack can never leave a half CLI that looks installed.

A pin is only used when it is newer than the version bundled in the app. The engine files for Codex, Gemini and the Claude Code CLI lane each check for a pinned binary first, and fall back to the bundled one. The comment in the Claude Code lane says what it is for: “a model released after this build shipped runs without waiting for an app update.” Older versions are pruned, and a bad download is deleted while the bundled CLI keeps running.

07What Opus 5.5 taught us

A row alone would have shipped a dead model.

Opus 5.5 is exactly the case the system was built for, and it still needed a release. Here is why.

Luminair has two ways to run Claude. The built-in lane runs turns inside the app through Anthropic's Agent SDK. The Claude Code CLI lane spawns the claude program like any other engine. The engine file for the built-in lane says so plainly: it is run “through the Agent SDK rather than by spawning a CLI”, and main.js hands the SDK the Claude binary bundled with the app. There is no pin lookup on that path.

The first problem was the version gate, not the model. Our project notes from 23 September record it: the API refuses Opus 5.5 below Claude Code 2.1.280. The SDK we bundled, 0.3.258, carried an older CLI. A catalog row would have put Opus 5.5 in the picker, and every turn on the built-in lane would have been refused.

So it shipped the old way, fast. Luminair 1.0.183, committed on 23 September, the day after launch, moved the bundled Agent SDK to 0.3.281, which carries Claude Code 2.1.281. It added Opus 5.5 to both Claude lanes, the price table ($4 in, $20 out, $5 cache write, $0.20 cache read per million) and the model's role card, whose text now ends: “Needs Claude Code 2.1.280 or newer, which this build bundles.” It was verified live with a one-line prompt.

The picker is a promise. A model in it that cannot answer is worse than a model that is not there yet.

The lesson we took is procedural. When a new Claude model appears, try it on the bundled CLI with a one-line prompt before adding any row. If the vendor says the tool is too old, move the SDK in the same pass. The catalog is still the fast path for Codex and Gemini builds, for the CLI lane, and for rows that only need a label.

One more thing from the launch coverage is worth knowing if you run agents. The New Stack points out that when Opus 5.5's safety classifiers fire, “Anthropic reroutes the request transparently to an older model.” That happens on Anthropic's side. No catalog on your machine can see it.

08What we checked

Checked, and not claimed.

7/7
Catalog tests pass: signature, rollout gating, per-login blocks, CLI pins, discovery, canary, and the hash-checked installer
3
Rollout stages (10, 50 and 100 percent), at least an hour apart, advanced by the backend watcher
>50%
Failure rate over at least 20 turns in two hours that sets the kill switch
1 day
From Opus 5.5's launch to the release that could run it on the built-in lane

What this post does not claim

  • That every new model reaches you without an app update. The built-in Claude lane, the price table and the role cards still ship with the app.
  • That a CLI pin was used for any particular model. This post describes the mechanism in the code, not a log of which rows or pins were served.
  • How fast any given install picks up a row. It depends on the hourly refresh and on the rollout stage.

For how Luminair chooses among all these models on each turn, read The right model for every turn.

Sources

New models, as they land

Open the model picker and see what arrived since your last update.

Download Luminair →