The short version
- MCP lets an agent use outside tools. Most clients load every tool's description up front. Anthropic has seen tool definitions consume 134K tokens before any work begins.
- The fix in the industry is the same idea twice: load less, later. Tool search loads definitions on demand; Agent Skills show only a name and a description until a skill is needed.
- Luminair keeps one list of your connected MCP servers and translates it into each engine's own format: SDK options for Claude, config overrides for Codex, a settings file for Gemini.
- Our own worst context bloat was not a tool. A 73KB reference document in the always-on Harness folder took per-turn prompts from about 30KB to about 100KB. Harness notes over 24KB, and all docs, are now left out.
Paying before you ask.
The Model Context Protocol, MCP, is a standard way to plug tools into an AI agent: a browser, a design app, a database, your calendar. Each server describes its tools in plain text, so the model knows what exists and how to call it. That description has to live somewhere, and by default it lives in the model's context window, the same limited space as your code and conversation.
Anthropic's engineering team wrote this down plainly in November 2025. In Code execution with MCP: “Most MCP clients load all tool definitions upfront directly into context”. With enough servers, “they’ll need to process hundreds of thousands of tokens before reading a request.” Their example workflow, rewritten so the agent writes code against the tools instead of calling each one directly, went “from 150,000 tokens to 2,000 tokens”. That 150,000 is a whole task's cost, results included, not the size of a list.
Three weeks later, Introducing advanced tool use put numbers on the list itself. Five common servers came to “58 tools consuming approximately 55K tokens before the conversation even starts.” And: “At Anthropic, we've seen tool definitions consume 134K tokens before optimization.” Their tool search approach loads definitions only when the model looks for them; for their example they report an “85% reduction in token usage”.
A large context window does not make this free. Every token of tool text is read on every turn, costs money or quota, and competes for the model's attention with the thing you actually asked.
A table of contents, not the book.
Agent Skills, which Anthropic introduced in October 2025, apply the same idea to instructions. A skill is a folder with a SKILL.md file and whatever scripts or references it needs. Anthropic's engineering post calls the principle out by name: “Progressive disclosure is the core design principle that makes Agent Skills flexible and scalable.” In practice: “At startup, the agent pre-loads the name and description of every installed skill into its system prompt.” The body loads only when the task calls for it.
Luminair uses skills in two ways. It ships a small bundled plugin with its own default skills, so commands like //atlas and //sonar work on any Mac without setup. And it can install a skill from a Git repository into ~/.claude/skills, the folder the Claude Agent SDK already scans, after which it is available in every Claude session.
Connectors are not Claude-only.
Luminair runs many engines, and each speaks MCP in its own dialect, if at all. We did not want a connector that works on Claude and silently vanishes when you switch a session to Codex. So there is one registry, MCP_PRESETS in main.js, and one function, extraMcpServers(), that returns every server that could actually be dialled right now: the key is saved, the binary exists, the URL is set.
The presets are data, not plumbing. Each one says how it runs (a local command or a remote URL), what it needs (a key, a URL, both, or nothing) and, for local servers, which environment variable carries the key. Today the list includes Playwright for browser control, Cloudflare, Composio, ElevenLabs, Unity, stock photo search and a few bring-your-own-URL entries. Where no official server exists, the code says so and asks you for the URL rather than inventing one.
The shared runner then asks each engine to translate. An engine file may declare a cli.mcp hook; if it does, the runner hands it the list and merges back whatever flags, environment or config it returns. If it does not, the engine is untouched.
Codex: flags, never secrets
Codex accepts servers as -c mcp_servers.<name>.… config overrides. The catch is that a command line is readable by every process running as you on a Mac. So the Codex hook never puts a secret value in the arguments. Keys and tokens travel as environment variables, and the config only names which variables to read. The code is equally honest about a limit: Codex's HTTP transport only supports a bearer token, so a server that authenticates with a custom header, such as Composio's x-api-key, “is SKIPPED rather than sent with auth Codex would drop and then fail on confusingly.”
Gemini: a settings file per account
The Gemini command line tool has no per-run MCP flag. It reads mcpServers from settings.json in its home folder, and each Gemini account in Luminair already has its own home. So the hook rewrites just that key before the turn, keeps everything else in the file, and writes secrets as references like ${OW_MCP_…_TOKEN} whose values arrive through the environment. The file is written readable by you alone. Gemini supports arbitrary headers, so every server carries over.
Claude: out of the process list
The Claude Agent SDK turns its MCP options into a --mcp-config argument on the command line it spawns. That would expose connector keys the same way. Luminair gives the Claude process only a pointer on the command line and sends the real JSON over an inherited pipe (a short-lived private file on Windows).
Finally, every engine is told what is wired in. The session contract carries a short block, “CONNECTED MCP SERVERS”, that lists the servers and tells the model connectors are not Claude-only.
Our biggest bloat was a note.
Luminair has a Journal, a folder of Markdown notes it can read and write. One subfolder, Harness, is special: its notes ride along on every turn of every session. That is right for short rules; on our own machine the folder holds the app rules and a note on how to check evidence. It turned out to be wrong for everything else.
On 10 September 2026, a 73KB reference document landed in that folder. From then on, about 70KB of every prompt, in every session, was that document. The code comment that records the fix is specific: “Per-turn prompts went from ~30KB to ~100KB, transcripts grew 3x faster”, and the large resumed transcripts that followed pushed Claude command line turns past the start watchdog. The comment is dated 14 September, four days after the document landed.
The fix is small on purpose. Docs are for looking up, so a model searches the Journal for them when it needs them. That is progressive disclosure again, arrived at by accident.
The same instinct shows up elsewhere. On resumed Antigravity threads, Luminair stops re-sending the static parts of its own session rules, including the connected-servers block, because the engine's transcript already holds them; we measured 44% of one 3.6 MB transcript as repeated boilerplate.
Deferred, except when it hides.
On the Claude lane, tool schemas load on demand. The context meter says so in its own words: “an attached server costs a name, not its whole schema, until it is used.”
But deferral has a failure mode, and we hit it. On 30 August a session “can't see the connector”: its tools sat behind search and the model never looked. So a few servers are marked to load straight into the prompt instead: the stock photo servers, Openverse, Wikimedia Commons and ElevenLabs. Everything else stays deferred. It is a judgement call per server, and we would rather make it visibly in one registry than leave it to chance.
See it, then trim it.
- 1Type //context in any session. The sheet lists the blocks Luminair injects and, under Tools attached (MCP), only the servers that are connected.
- 2Untick Include connected tools in the same sheet to run the next turns without them, and compare.
- 3Keep Harness notes short. If a note is reference material, save it as a doc in another folder; the model can still search for it.
Checked, and not claimed.
What this post does not claim
- How many tokens your own connected servers cost. It depends on the servers, and our meter does not count schemas.
- That the Anthropic figures apply to Luminair sessions. They are Anthropic's measurements of their own setups.
- That every engine can use every connector. Codex skips custom-header servers, and engines without an MCP hook get none.
Sources
- Anthropic Engineering · 4 November 2025Code execution with MCP: building more efficient AI agents
- Anthropic Engineering · 24 November 2025Introducing advanced tool use on the Claude Developer Platform
- Anthropic Engineering · 16 October 2025Equipping agents for the real world with Agent Skills
For more on what fills a model's context, read Your model didn't get dumber.
Connect once, use everywhere
Plug in a server and every engine that speaks MCP gets it, without its keys on a command line.