← Blog Systems & Memory 26 September 2026 9 min read

Context rot, from the client side

Researchers have shown that models get less reliable as their input grows. From inside an app that resumes long sessions all day, there is a second problem nobody benchmarks: long context is slow, and slow looks a lot like broken.

Fig 01  What a long session sends on every turnSchematic · not to scale
ONE RESUMED TURN · ONE REQUEST Fixed overhead system prompt, hooks, tool schemas, skills oldest turns trimmed past the cap Transcript every earlier message, tool call and tool result, resent Turn contract Your message 0 TOKENS A 0.8 MB TRANSCRIPT WAS ABOUT 250K 1 · SLOW Silence before the first byte A 60 second watchdog read it as a hang. Fix: past init, a 5 minute grace 2 · MIS-MEASURED A cap sized as bytes ÷ 4 Real rates ran 2.0 to 6.6 bytes a token. Fix: the tokens the API reported 3 · FORGOTTEN Compaction drops what mattered Including a draft written three turns ago. Fix: keep it outside the window
The proportions are illustrative. The 0.8 MB and roughly 250k-token figures come from a comment in desktop/engines/runner-cli.js; the 2.0 to 6.6 range is from the commit that replaced the byte guess.

The short version

  1. “Context rot” is the finding that models handle long inputs less reliably than short ones, even on simple tasks. Bigger windows did not make it go away.
  2. From the app side, long context also means long waits. A resumed session of 0.8 MB is a request of roughly 250,000 tokens, and the model can be silent for over a minute before it answers.
  3. Luminair used to kill those turns as hung. It now waits up to five minutes once the engine has proved it started.
  4. The per-session size cap now counts the tokens the API reports instead of guessing from bytes, and a turn contract kept outside the window carries what compaction would lose, now including working drafts.
02Built with Luminair

Our sessions are long.

Luminair is written in Luminair, and our own sessions are the long kind: days of work on one feature, resumed again and again, full of tool output. That makes us the first to hit every long-context failure the app has. Each fix below started as one of our sessions going wrong, and each is recorded in the code with a date and the exact error the user saw.

This post was written the same way, in a Luminair session reading the code and the commit history, with outside claims checked against the pages linked at the end.

03Longer is not better

The window grew. The reliability did not.

In July 2025, Chroma published Context Rot: How Increasing Input Tokens Impacts LLM Performance. They tested 18 models on tasks kept deliberately simple, changing only the length of the input. The result: “models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows.”

The report also explains why this surprised people. The usual long-context test hides a fact in a long document and asks for it back, and models score close to perfect on it. But that test mostly checks exact word matching. Real work, like an agent picking through a long session, asks for more. Chroma's conclusion is a line every harness builder should keep in view: “Whether relevant information is present in a model’s context is not all that matters; what matters more is how that information is presented.”

The labs have been working on it. When Anthropic released Claude Opus 4.6, InfoQ reported a 1M token context window in beta and called context compaction “the more significant architectural update”. It “addresses performance degradation as context windows fill, a phenomenon Anthropic calls ‘context rot.’” When a conversation nears the limit, the API “automatically summarizes earlier portions and replaces them with a compressed state.”

Compaction helps. It also means the model is now working from a summary, and a summary can leave out the one thing you needed. More on that in section 06.

A bigger window is a bigger desk. It does not make the reader any faster.
04Longer is slower

Silence is not a hang.

Research on context rot measures answer quality. What an app notices first is time. Every resumed turn sends the whole conversation again, and the model has to read all of it before it writes a word. On a long session, that first word can take more than a minute.

Luminair runs command-line engines, including Claude's, as child processes. Each turn has a start watchdog: if nothing useful comes back in 60 seconds, the app assumes the tool failed to start, kills it and says so. That rule exists for a good reason. A tool that never boots should not leave you staring at a spinner forever.

On 3 September a user reported the loop from inside a session: “Claude started but produced no usable output for 60 seconds”, then “continue”, then the same error. The first fix treated a status line the Claude tool prints when it sends its request as proof of life. On 14 September the same error came back, twice on one resumed session.

This time we reproduced it. Pointing the bundled Claude tool (version 2.1.258) at a stalled dummy server showed exactly what happens:

Fig 02  What the Claude tool prints on a long resumed turnFrom the reproduction notes
0 sprocess startsHooks run, tool servers connect.
~2 ssystem / initThe tool has booted and accepted the prompt. It sends the request.
silencenothing at all on the streamThe model is reading about 250k tokens.
60 sold rule: killed as “no usable output”
laterfirst response bytesOnly now does the “requesting” status line appear, with the answer.
Reconstructed from the comments in desktop/engines/runner-cli.js and desktop/test/core/start-watchdog.test.js. The ~2 s and the 250k-token size are from those notes; how long the silence lasts depends on the model and the session.

The status line the first fix relied on turned out to arrive only once the server starts streaming, which is after the silence, not during it. The real signal was already there: the init event. By the time it appears, the tool has booted, connected its tool servers, run its hooks and accepted the prompt. That is everything the 60 second watchdog was ever meant to guard.

So now, when init arrives, the runner trades the 60 second watchdog for a single five minute grace. If the answer starts in that window, you get it. If five minutes pass in silence, the turn is still stopped, with a message that says five minutes, not sixty seconds. A tool that never reaches init is still killed at 60 seconds, as before. Codex got the same treatment from its own start signal a week earlier; we told that story in Your model didn't get dumber.

Why this is context rot tooThe user saw an error, typed “continue”, and resent the same long context, which failed the same way. A slow model looked like a broken one. The size of the context was the cause in both cases; only the symptom differed.
05Count real tokens

Four bytes a token was a guess.

Every Luminair session has a maximum size. You pick it per session, from 1k to 1M tokens or no limit, and new windows start at 100k. When a session grows past its cap, the oldest turns leave the live transcript. They are archived, not deleted, and by default they leave a short summary behind.

The menu always spoke in tokens. The code did not. Until 7 August 2026 it enforced the cap as the token number times four bytes, a common rule of thumb that, in the commit's words, “was never calibrated against anything.” Measured on real transcripts, the true rate ran from 2.0 to 6.6 bytes a token, so the old guess was off by up to 65% either way. Tool-heavy sessions full of JSON pack very differently from plain chat.

The fix uses numbers that were already on disk. Every assistant record in a Claude transcript carries the usage the API reported for that turn. Between two such turns, the prompt grew by some number of tokens, and the lines written in between are what caused it. Divide one by the other and you have this session's own rate.

Fig 03  Measuring a session's own bytes per tokenSchematic
TRANSCRIPT, OLDEST TO NEWEST usage: P1 you tool call tool result usage: P2 B bytes of transcript that caused the growth rate = Σ B ÷ Σ (P2 − P1) summed over every pair of turns where the prompt grew 0 2 measured on real transcripts: 2.0 to 6.6 old flat 4 clamp 12
The method from contentBytesPerToken in desktop/main.js and bytesPerTokenFrom in desktop/lib/core/trim-plan.js. The 2.0 to 6.6 range is from the 7 August 2026 commit. With nothing to measure yet, the rate falls back to 4.

There is one subtlety, and the code spells it out. Most of a live context is not transcript at all: system prompt, hooks, tool schemas, skill docs. The commit notes that the session it was written in read 379k tokens against a 213 KB file. Trimming history cannot remove any of that overhead, so a cap measured against the whole context “would trim forever and never get under.” The cap is measured against the transcript's own share.

Counting correctly is not enough if you count the wrong file. On 26 September we found the cap had stopped working for sessions run through the Claude command-line tool: the trimmer looked in one store while the tool resumed from another. A session set to 25k had grown to 205k tokens. It now trims every copy it can find.

06Keep state outside

What compaction cannot touch.

Engines own their compaction. Luminair cannot stop Claude or Codex from summarising an old stretch of conversation, and a summary is lossy by design. What the app can do is keep a small, structured record outside the window and put it back on every turn.

That record is the turn contract, shipped on 29 August 2026 as part of Harness v1. The code calls lossy compaction “the #1 harness complaint in the field”. The contract is rebuilt from durable state, never from a model: the prompts you actually sent, the recent turns, the files changed on disk (from snapshots taken before each turn, so shell-made changes count too) and the files already read. It is capped at 3,400 characters, because, as the source puts it, “the contract must never crowd the window it protects.”

Injected on every turn · turn contract
=== TURN CONTRACT (rebuilt by the app OUTSIDE the context window: survives any compaction) ===
If this conversation was compacted or feels incomplete, TRUST THIS BLOCK over your memory of the thread.
SESSION STARTED WITH:
LATEST ASK:
TURNS SO FAR: . Recent:
  -  
ACTIVE WORKING DRAFTS & DELIVERABLES (preserved outside context: survives any compaction):
[Draft: "" · words · Turn ]
FILES THIS SESSION ACTUALLY CHANGED ON DISK (from pre-turn snapshots: includes shell-made changes):
  -
FILES THIS SESSION ALREADY READ (from the tool stream; ...)
  -
=== END TURN CONTRACT ===
Headings copied from desktop/lib/compaction-contract.js; contents left blank on purpose. The ruled section is the one added on 25 September 2026. The same record is mirrored into the session's Journal doc under Short Term Memory, so you can read it too.

The draft that disappeared

The contract had a blind spot. It tracked files and prompts, not what the assistant had written in chat. For coding that is mostly fine: the work is on disk. For writing it is not. A letter or an email drafted over several turns lives only in the conversation.

On 25 September a commit titled “Fix draft loss from research compaction” closed that gap. A research report had been written into a session's transcript in full, the context filled, and compaction took the earlier turns with it, including a working draft. The fix has three parts:

01

Drafts are kept outside the window

After each finished turn, the app looks for a deliverable in the reply: a Subject line, a heading like “Revised Email”, a letter with a greeting and a sign-off, or a prompt that asked for a draft. It keeps up to five per session, updates a revision in place instead of duplicating it, and puts them in the contract.

02

They are saved as Journal notes

Each detected draft is also written to the Journal as a note in a Drafts folder, so it exists as a file, not only as chat text.

03

Big research stays out of the transcript

A research result over 8,000 characters is cut to 7,500 in the chat, with a line pointing to the full report, which is saved as a Journal note.

The same commit added a rule to the prompt: a model must not say a draft “was never provided” or “was made up” because it cannot see it. It should assume compaction happened and check the contract, the session doc and the transcript first. Forgetting is sometimes unavoidable. Denying is not.

07Find it in the app

Three settings, one minute.

  1. 1In a session, open the ⋯ menu and choose Max size to set that session's cap. The default for new windows is under Settings › Sessions › Window token limit.
  2. 2In the Solace panel, Summarize aged-out history decides whether trimmed turns leave a short summary behind. It is on by default.
  3. 3Open the session's Journal doc and read Short Term Memory and Active Drafts: the same state the model gets after a compaction.
08What we checked

Checked, and not claimed.

3/3
Start-watchdog tests pass: silent request survives, dead start is killed, silence past the grace is killed
6/6
Compaction draft tests pass, including revisions and branching a session
9/9
Trim planner and trim retry tests pass
3,400
Character cap on the turn contract, raised by 5,000 when it carries drafts

What this post does not claim

  • That a smaller cap gives better answers. Chroma's results suggest shorter input helps; we have not measured it on Luminair sessions.
  • How long the silence before a first byte usually lasts. The one measurement we cite is the 60 second threshold it crossed.
  • That draft detection catches every deliverable. It looks for common shapes: subject lines, headings, letters and explicit requests.
  • That the byte-rate method applies to every engine. It reads the usage records in Claude transcripts.

Sources

For how the Journal stores memory without a vector database, read Memory without a vector database.

Keep long sessions healthy

Set a cap per session, and let the contract carry what compaction drops.

Download Luminair →