The short version
- A model is like a car engine. The harness is the rest of the car: the gearbox, the fuel line, the dashboard. When the ride gets rough, the engine is the first suspect and often the wrong one.
- In Anthropic's two public postmortems, the causes were serving bugs (2025) and harness changes in Claude Code (2026). Neither blamed the weights.
- We found the same pattern in our own code: a watchdog that killed slow-thinking Codex turns, and an SDK option that silently dropped our system-prompt additions.
- Luminair gives you two checks: //context shows what is in the window, and switching model mid-thread gets a second opinion with the conversation carried over.
The bugs came from using it.
Luminair is built inside Luminair. The two bugs in this post were not found by a test suite. They were found by running real work through the app, on Codex and on Claude, and noticing that something felt off. Both fixes were written in Luminair sessions and carry a Claude co-author line in git.
This post was drafted the same way: one session reading the repository and its history directly, with every outside claim checked against the source it links to.
“It got worse.”
The feeling is familiar. Last week the assistant fixed a bug in one go. This week it forgets what it was doing, repeats itself, or gives up early. The easy explanation is that the provider quietly swapped in a cheaper model. The New Stack ran that theory under the headline “AI shrinkflation: Why Anthropic's Claude Opus 4.7 may be less capable than the model it replaced” on 23 April 2026.
The same day, Anthropic published an update on recent Claude Code quality reports. It named three causes, and none of them was the model:
A default changed
Claude Code's default reasoning effort moved from high to medium, to cut long waits. Anthropic called it “the wrong tradeoff” and reverted it on 7 April.
Memory got wiped, over and over
A change meant to clear old thinking once, after an idle hour, cleared it on every turn instead. That “made Claude seem forgetful and repetitive”. Fixed 10 April. The postmortem adds that the resulting cache misses likely drove separate reports of usage limits draining faster than expected.
A line in the system prompt
An instruction to keep text between tool calls to 25 words or fewer, and final answers to 100, “hurt coding quality”. Reverted 20 April.
The post is explicit: “We never intentionally degrade our models, and we were able to immediately confirm that our API and inference layer were unaffected.” And because each change hit a different slice of traffic on a different schedule, “the aggregate effect looked like broad, inconsistent degradation.”
This had happened before. In September 2025, Anthropic's postmortem of three recent issues traced a wave of complaints to three infrastructure bugs: requests routed to the wrong server pool, a misconfiguration that corrupted output, and a compiler bug in token selection. At the worst hour, 16% of Sonnet 4 requests were misrouted. Its plainest line: “We never reduce model quality due to demand, time of day, or server load.”
Our harness did it too.
Luminair is a harness. It wraps Claude, Codex, Gemini and a dozen other engines in one app, which means it adds its own layer of prompts, timeouts and retries on top of theirs. Two recent fixes show how easily that layer can make a strong model look weak.
1 · The watchdog that punished thinking
Every engine turn in Luminair has a start watchdog: if a command-line engine prints nothing useful for 60 seconds, the app assumes it hung, kills it and says so. That is reasonable for a tool that failed to boot. It is wrong for a model that is thinking hard.
On the morning of 7 September, six Codex turns died with “produced no usable output”. The session was large (a resumed thread of about 5 MB, on GPT-6 Astra) and the model took longer than a minute before its first reply. The same prompts worked whenever the model happened to answer faster. From the user's seat, it looked exactly like a model that had got flaky.
The fix was to listen more carefully. The Codex command-line tool prints a turn.started line the moment the prompt is on its way to the model. Luminair used to ignore it. Now the Codex engine file turns it into a quiet “alive” signal:
// Proof of life. The CLI prints turn.started as soon as the prompt is on its way to // the model, but the first item.completed (reasoning or message) can take minutes on // a large resumed thread. Without this the runner's 60 s start watchdog killed such // turns as "produced no usable output" (2026-09-07, six times in one morning). if (ev.type === 'turn.started') return { t: 'alive' };
The shared runner treats “alive” as proof the prompt reached the model and swaps the 60 second watchdog for a single five minute grace. It is never shown to you and never counts as progress. In mid-September the Claude lane got the same treatment: once Claude's command-line tool reports it has started up, a long silence before the first word is no longer mistaken for a hang. A test now checks all three cases: a silent request that survives, a tool that never starts and is still killed, and a silence past the grace that is killed with the five minute message.
2 · The system prompt that went nowhere
The second bug is closer to Anthropic's April story. Luminair adds its own instructions to Claude turns: app rules, your abilities, the output format our verifier expects. We passed them to the Claude Agent SDK as an option called appendSystemPrompt.
On 27 September we found that the SDK version we bundle, 0.3.281, never reads that option. Its own normalizer only looks at systemPrompt, and when that is missing the system prompt is empty. So those rules reached no model, and the query also ran without Claude Code's standard preset. We had raised the dependency to that version on 23 September; we did not check whether older versions read the option. The fix folds the old option into the shape the SDK actually reads, and a probe confirmed the same text is ignored one way and obeyed the other.
Nothing crashed. There was no error to see. A model that has lost its instructions simply behaves a little worse, which is exactly the kind of drift people blame on the model.
What is the model actually reading?
Think of the model's context window as a desk. Before your question lands on it, the harness has already put papers there: house rules, notes from earlier, the tools it may use. If the desk is cluttered, or the one page that mattered never made it, the answer suffers.
Harness v2, which landed on 6 September, added a context-budget meter. Type //context in any session and a sheet lists every block Luminair injects, what each costs in tokens, which tool servers are attached, and what the last turn actually processed. After a turn has been sent, it switches to the blocks captured at that dispatch, so you see what went out rather than a guess.
Context budget · gpt-6-astra
Tokens are estimated at four characters each. These blocks were captured at the last dispatch, . Luminair dispatch blocks only; engine-owned history, native tools, skills and later hooks are not measured.
Context tree: own tokens against the sessions this one handed work to.
Include optional Journal notes and old conversation summaries Include connected toolsTwo details matter. First, the sheet says what it cannot see: history the engine keeps for itself, its native tools and later hooks. A meter that overclaims is worse than none. Second, the two checkboxes let you drop optional Journal notes, old summaries and connected tools from the next turns, so you can test whether a lighter desk gives a better answer.
A second opinion, mid-thread.
The fastest way to tell a model problem from a harness problem is to change one thing. If a turn feels wrong, send the next one to a different model, in the same session, with the same history. If the other model does fine, you learned something. If both stumble, look at the harness, or the task.
In Luminair you can switch a session's model at any point, including to another company's engine. The other engine has never seen this conversation, so Luminair writes it a handoff: a catch-up note placed before your next message that says, in effect, “this session ran on another model; here is the conversation so far; carry on”.
buildEngineBridge in desktop/main.js packs the handoff. It walks back from the newest turn, keeps whole turns until the budget is spent, and tells the model how many earlier turns it left out. Bar lengths are illustrative.That budget is a real trade-off, and it is worth being honest about. A fixed catch-up note is cheaper and faster than replaying everything, but it means the new model sees a trimmed, text-only version of a long thread. For a second opinion on the last few turns, that is usually what you want. For a thread with a hundred turns of subtle context, expect the new model to know less than the old one did.
The bridge is also visible. When a turn carries one, the //context sheet lists it as Cross-model history bridge with its size. And once the new engine has answered, the turn is written back into the session's own record, so switching back later only has to catch up on what the first model missed.
Three checks, two minutes.
- 1In the session that feels off, type //context in the composer and press Enter. Look for a block that is huge, missing, or not what you expected.
- 2Click the model name in the session's header (or open the session's ⋯ menu and choose Model) and pick a model from another engine. Resend or ask again. Models on engines other than Claude need Pro.
- 3Type //doctor, or open your avatar at the bottom left, then Settings › General › Harness health › Check now. It checks logins, engine binaries, connectors and more, with one-tap safe repairs.
Every // command is listed under Settings › General › Luminair Abilities › View.
Checked, and not claimed.
What this post does not claim
- That no model has ever changed. We can only speak to the causes the providers published and the bugs in our own code.
- How much better answers got after either fix. We saw the failures stop; we did not measure answer quality.
- That a second model's answer is proof. It is a signal, and a cheap one.
If you want the longer story of how Luminair picks a model per turn, read The right model for every turn.
Check your own harness
Open any session, type //context, and see what your model is really reading.