The short version
- A new model needs two things on your machine: a row in the model picker, and a runner (the vendor's command-line tool) new enough that the vendor will accept it.
- Luminair fetches a model catalog every hour. It is signed with Ed25519, rolled out to 10, 50 and then 100 percent of installs, and has a kill switch per row.
- The same catalog can pin a newer vendor CLI. The app downloads it from the npm registry, checks its sha512 hash before unpacking, and prefers it over the bundled copy.
- Opus 5.5 needed Claude Code 2.1.280 or newer. Our built-in Claude lane runs the CLI bundled in the app, so that one took a release: Luminair 1.0.183, the day after launch.
A new model every few weeks.
On 22 September 2026, Anthropic introduced Claude Opus 5.5. Its announcement says it “performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.” Input and output are $4 and $20 per million tokens. TechCrunch noted that the launch “comes just two months after the release of Opus 5 on July 24.”
Our own git history shows the same rhythm from the app side. Opus 5 was added to Luminair's Claude engine on 29 July, Fable 5.1 on 1 September, and Opus 5.5 on 23 September. On the OpenAI side, GPT-6 Astra landed in early September. That is four new top models in the picker inside eight weeks, and each one is a moment where users ask, reasonably, “why can't I pick it yet?”
If every new model waits for an app build, then signing, notarizing, uploading and waiting for every Mac to install the update, you are always a few days behind. So on 7 September we built a way around it. The file that holds the logic opens with the goal in plain words: “a new model appears the moment it is released, with no app update”.
The picker is the easy half.
Think of a new model like a new TV channel. You need it in the channel guide, and you need a receiver that can decode it. Put it in the guide without the decoder and you get a blank screen with a nice label.
Luminair runs most engines by driving the vendor's own command-line tool: Codex for OpenAI, Gemini CLI for Google, Claude Code for Anthropic. Vendors can gate a new model on the version of that tool. GPT-6 Astra is where that bit us. The comment on its row in desktop/engines/codex.js records it: the Codex CLI we bundled, 0.147, “got ‘requires a newer version of Codex’ from OpenAI”. It ran only with 0.153 or newer, verified on 7 September with 0.153.4.
That is why the catalog has two kinds of entry. A model row carries only what the picker needs: an id, a label, a short name, a colour and a rank. A CLI pin names an exact build of the vendor's tool, per platform. Both are optional, and a new model often needs both.
Anything that isn't signed is ignored.
A remote list that can add models, and download programs, is a tempting target. So the catalog is signed. The backend holds an Ed25519 private key; the app ships with the public half baked into its code. The served body is a small envelope: a version, a key id, the catalog as a string, and a signature over that exact string.
// Signing the exact string (not a re-serialized object) keeps verification // independent of key order and number formatting. const key = crypto.createPublicKey(publicKeyPem); ok = crypto.verify(null, Buffer.from(body.payload, 'utf8'), key, Buffer.from(body.sig, 'base64')); if (!ok) return null;
Signing the string rather than the parsed object is a small choice that avoids a class of bugs: two JSON encoders can disagree about key order or how to write a number, and then an honest catalog fails to verify.
The app fetches the catalog three seconds after launch and then every hour. The last good body is cached on disk and re-verified on every load, so the first engine list at boot is already right, even offline, and a tampered cache is treated exactly like a tampered download: ignored. Because the public key lives in the code, rotating it means a new app build. That is the price of trusting nothing else.
Ten percent, then fifty, then everyone.
A signed row is trusted, but not yet proven. Each row can carry a rollout percentage. Every install hashes its device id together with the model's key into a number from 0 to 99, and sees the row only if that number is under the rollout. Because the bucket never changes for a given install and model, raising the number only ever adds people. Nobody sees a model appear, vanish and reappear between refreshes.
The backend's hourly watcher moves a healthy row one stage at a time, 10, 50, 100, waiting at least an hour per stage. It reads anonymous telemetry: per model, a hashed install id, the outcome, and no prompt or account. If at least 20 turns ran in the last two hours and more than half failed, it sets kill and raises an alert that the row was rolled back. A killed row disappears for everyone, admins included. CLI pins go through exactly the same gate.
A model the watcher discovers for a command-line engine has to pass a canary on the admin's Mac, which has real logins, before it is promoted to 10 percent: a chat reply and a tool call, each asked to echo a random token like OWCANARY-…. Passing is a string match, not a judgment. Only chat plus tool on at least one login counts as a pass.
Downloading a program, carefully.
A CLI pin says: for this engine, on this platform, use this npm tarball, with this sha512 integrity hash, and the program lives at this path inside it. desktop/lib/cli-updater.js then does the install in a way that fails closed.
Only the registry
A pin whose tarball URL is not on registry.npmjs.org is refused, both when the pin is read and again before download.
Hash first
The bytes are checked against the pinned sha512 before anything is unpacked. A mismatch throws and nothing is written outside a scratch folder.
Files only
Unpacking keeps regular files and folders only: no symlinks, no absolute paths, no paths climbing out with “..”.
Swap in whole
Everything happens in a .partial folder, then one rename moves it to cli/<engine>/<version>. A crash mid-unpack can never leave a half CLI that looks installed.
A pin is only used when it is newer than the version bundled in the app. The engine files for Codex, Gemini and the Claude Code CLI lane each check for a pinned binary first, and fall back to the bundled one. The comment in the Claude Code lane says what it is for: “a model released after this build shipped runs without waiting for an app update.” Older versions are pruned, and a bad download is deleted while the bundled CLI keeps running.
A row alone would have shipped a dead model.
Opus 5.5 is exactly the case the system was built for, and it still needed a release. Here is why.
Luminair has two ways to run Claude. The built-in lane runs turns inside the app through Anthropic's Agent SDK. The Claude Code CLI lane spawns the claude program like any other engine. The engine file for the built-in lane says so plainly: it is run “through the Agent SDK rather than by spawning a CLI”, and main.js hands the SDK the Claude binary bundled with the app. There is no pin lookup on that path.
The first problem was the version gate, not the model. Our project notes from 23 September record it: the API refuses Opus 5.5 below Claude Code 2.1.280. The SDK we bundled, 0.3.258, carried an older CLI. A catalog row would have put Opus 5.5 in the picker, and every turn on the built-in lane would have been refused.
So it shipped the old way, fast. Luminair 1.0.183, committed on 23 September, the day after launch, moved the bundled Agent SDK to 0.3.281, which carries Claude Code 2.1.281. It added Opus 5.5 to both Claude lanes, the price table ($4 in, $20 out, $5 cache write, $0.20 cache read per million) and the model's role card, whose text now ends: “Needs Claude Code 2.1.280 or newer, which this build bundles.” It was verified live with a one-line prompt.
The lesson we took is procedural. When a new Claude model appears, try it on the bundled CLI with a one-line prompt before adding any row. If the vendor says the tool is too old, move the SDK in the same pass. The catalog is still the fast path for Codex and Gemini builds, for the CLI lane, and for rows that only need a label.
One more thing from the launch coverage is worth knowing if you run agents. The New Stack points out that when Opus 5.5's safety classifiers fire, “Anthropic reroutes the request transparently to an older model.” That happens on Anthropic's side. No catalog on your machine can see it.
Checked, and not claimed.
What this post does not claim
- That every new model reaches you without an app update. The built-in Claude lane, the price table and the role cards still ship with the app.
- That a CLI pin was used for any particular model. This post describes the mechanism in the code, not a log of which rows or pins were served.
- How fast any given install picks up a row. It depends on the hourly refresh and on the rollout stage.
For how Luminair chooses among all these models on each turn, read The right model for every turn.
Sources
- Anthropic · 22 September 2026Introducing Claude Opus 5.5
- TechCrunch · 22 September 2026Anthropic releases Opus 5.5 with lower prices and Fable-level performance
- The New Stack · 22 September 2026Anthropic releases Opus 5.5 and cuts pricing by 20%. Your agent calls might secretly get routed to an older model.
New models, as they land
Open the model picker and see what arrived since your last update.