← Blog Harness 14 September 2026 9 min read

Skills, MCP and the 134K-token tool list.

Every tool you connect to an agent costs context before you type a word. Anthropic measured it, then designed around it. Here is what that looks like from inside an app that hands the same tools to Claude, Codex and Gemini, and the one note that made all our prompts three times bigger.

Fig 01  One connector list, three dialectsSchematic · from the engine files
MAIN.JS extraMcpServers() Every connected preset, plus Google Docs, Buffer and Mobbin Only servers that could be dialled right now Claude lane SDK mcpServers, carried over a private pipe so no key shows in the process list ALL SERVERS Codex -c mcp_servers.<name>.… overrides; secrets passed by variable name, never value CUSTOM-HEADER SERVERS SKIPPED Gemini Rewrites mcpServers in the account's own settings.json, secrets as ${OW_MCP_…_TOKEN} ALL SERVERS Engines without an mcp hook are left untouched, so a connector can never break them
Drawn from extraMcpServers and secureSdkOptions in desktop/main.js, the cli.mcp hooks in engines/codex.js and engines/gemini.js, and the call site in engines/runner-cli.js. The Claude command line lane has its own hook too; it is folded into “Claude lane” here.

The short version

  1. MCP lets an agent use outside tools. Most clients load every tool's description up front. Anthropic has seen tool definitions consume 134K tokens before any work begins.
  2. The fix in the industry is the same idea twice: load less, later. Tool search loads definitions on demand; Agent Skills show only a name and a description until a skill is needed.
  3. Luminair keeps one list of your connected MCP servers and translates it into each engine's own format: SDK options for Claude, config overrides for Codex, a settings file for Gemini.
  4. Our own worst context bloat was not a tool. A 73KB reference document in the always-on Harness folder took per-turn prompts from about 30KB to about 100KB. Harness notes over 24KB, and all docs, are now left out.
02The tool list tax

Paying before you ask.

The Model Context Protocol, MCP, is a standard way to plug tools into an AI agent: a browser, a design app, a database, your calendar. Each server describes its tools in plain text, so the model knows what exists and how to call it. That description has to live somewhere, and by default it lives in the model's context window, the same limited space as your code and conversation.

Anthropic's engineering team wrote this down plainly in November 2025. In Code execution with MCP: “Most MCP clients load all tool definitions upfront directly into context”. With enough servers, “they’ll need to process hundreds of thousands of tokens before reading a request.” Their example workflow, rewritten so the agent writes code against the tools instead of calling each one directly, went “from 150,000 tokens to 2,000 tokens”. That 150,000 is a whole task's cost, results included, not the size of a list.

Three weeks later, Introducing advanced tool use put numbers on the list itself. Five common servers came to “58 tools consuming approximately 55K tokens before the conversation even starts.” And: “At Anthropic, we've seen tool definitions consume 134K tokens before optimization.” Their tool search approach loads definitions only when the model looks for them; for their example they report an “85% reduction in token usage”.

Fig 02  Published context costs of toolsReal figures · Anthropic, Nov 2025
Five servers58 tools55K
Anthropic, internalbefore optimization134K
One workflowdirect tool calls150K
Same workflowcode execution2K
075K150K tokens
The first two bars are tool definitions alone, from “Introducing advanced tool use”. The last two are a whole task's tokens, before and after, from “Code execution with MCP”. Different measurements, drawn on one scale for size only.

A large context window does not make this free. Every token of tool text is read on every turn, costs money or quota, and competes for the model's attention with the thing you actually asked.

03Skills load late

A table of contents, not the book.

Agent Skills, which Anthropic introduced in October 2025, apply the same idea to instructions. A skill is a folder with a SKILL.md file and whatever scripts or references it needs. Anthropic's engineering post calls the principle out by name: “Progressive disclosure is the core design principle that makes Agent Skills flexible and scalable.” In practice: “At startup, the agent pre-loads the name and description of every installed skill into its system prompt.” The body loads only when the task calls for it.

Luminair uses skills in two ways. It ships a small bundled plugin with its own default skills, so commands like //atlas and //sonar work on any Mac without setup. And it can install a skill from a Git repository into ~/.claude/skills, the folder the Claude Agent SDK already scans, after which it is available in every Claude session.

Show the model a table of contents. Hand it the chapter when it asks.
04One list, every engine

Connectors are not Claude-only.

Luminair runs many engines, and each speaks MCP in its own dialect, if at all. We did not want a connector that works on Claude and silently vanishes when you switch a session to Codex. So there is one registry, MCP_PRESETS in main.js, and one function, extraMcpServers(), that returns every server that could actually be dialled right now: the key is saved, the binary exists, the URL is set.

The presets are data, not plumbing. Each one says how it runs (a local command or a remote URL), what it needs (a key, a URL, both, or nothing) and, for local servers, which environment variable carries the key. Today the list includes Playwright for browser control, Cloudflare, Composio, ElevenLabs, Unity, stock photo search and a few bring-your-own-URL entries. Where no official server exists, the code says so and asks you for the URL rather than inventing one.

The shared runner then asks each engine to translate. An engine file may declare a cli.mcp hook; if it does, the runner hands it the list and merges back whatever flags, environment or config it returns. If it does not, the engine is untouched.

Codex: flags, never secrets

Codex accepts servers as -c mcp_servers.<name>.… config overrides. The catch is that a command line is readable by every process running as you on a Mac. So the Codex hook never puts a secret value in the arguments. Keys and tokens travel as environment variables, and the config only names which variables to read. The code is equally honest about a limit: Codex's HTTP transport only supports a bearer token, so a server that authenticates with a custom header, such as Composio's x-api-key, “is SKIPPED rather than sent with auth Codex would drop and then fail on confusingly.”

Gemini: a settings file per account

The Gemini command line tool has no per-run MCP flag. It reads mcpServers from settings.json in its home folder, and each Gemini account in Luminair already has its own home. So the hook rewrites just that key before the turn, keeps everything else in the file, and writes secrets as references like ${OW_MCP_…_TOKEN} whose values arrive through the environment. The file is written readable by you alone. Gemini supports arbitrary headers, so every server carries over.

Claude: out of the process list

The Claude Agent SDK turns its MCP options into a --mcp-config argument on the command line it spawns. That would expose connector keys the same way. Luminair gives the Claude process only a pointer on the command line and sends the real JSON over an inherited pipe (a short-lived private file on Windows).

Finally, every engine is told what is wired in. The session contract carries a short block, “CONNECTED MCP SERVERS”, that lists the servers and tells the model connectors are not Claude-only.

05Keep the desk clear

Our biggest bloat was a note.

Luminair has a Journal, a folder of Markdown notes it can read and write. One subfolder, Harness, is special: its notes ride along on every turn of every session. That is right for short rules; on our own machine the folder holds the app rules and a note on how to check evidence. It turned out to be wrong for everything else.

On 10 September 2026, a 73KB reference document landed in that folder. From then on, about 70KB of every prompt, in every session, was that document. The code comment that records the fix is specific: “Per-turn prompts went from ~30KB to ~100KB, transcripts grew 3x faster”, and the large resumed transcripts that followed pushed Claude command line turns past the start watchdog. The comment is dated 14 September, four days after the document landed.

Fig 03  What rides on every prompt nowRules from lib/vault/harness-filter.js
Before 10 Sep ~30KB per turn With the 73KB doc the same document, again, every turn ~100KB per turn · transcripts grew 3x faster
✓Rules and short notes ride on every prompt, frontmatter stripped.
✕Any note marked as a doc is left out, even a small one. Docs are reference, not rules.
✕Any note over 24KB is left out.
✕Empty notes are left out quietly.
Prompt sizes are the approximate figures recorded in the code comment, not a new measurement; bar lengths are proportional. The four rules are the whole filter, each pinned by a test (6 of 6 pass).

The fix is small on purpose. Docs are for looking up, so a model searches the Journal for them when it needs them. That is progressive disclosure again, arrived at by accident.

The same instinct shows up elsewhere. On resumed Antigravity threads, Luminair stops re-sending the static parts of its own session rules, including the connected-servers block, because the engine's transcript already holds them; we measured 44% of one 3.6 MB transcript as repeated boilerplate.

06The trade-off we made

Deferred, except when it hides.

On the Claude lane, tool schemas load on demand. The context meter says so in its own words: “an attached server costs a name, not its whole schema, until it is used.”

But deferral has a failure mode, and we hit it. On 30 August a session “can't see the connector”: its tools sat behind search and the model never looked. So a few servers are marked to load straight into the prompt instead: the stock photo servers, Openverse, Wikimedia Commons and ElevenLabs. Everything else stays deferred. It is a judgement call per server, and we would rather make it visibly in one registry than leave it to chance.

What we do not measureThe //context meter lists which MCP servers are attached, but it does not count their schema tokens, and it says so. Engine-owned history, native tools and skills are also outside what it measures.
07Find it in the app

See it, then trim it.

  1. 1Type //context in any session. The sheet lists the blocks Luminair injects and, under Tools attached (MCP), only the servers that are connected.
  2. 2Untick Include connected tools in the same sheet to run the next turns without them, and compare.
  3. 3Keep Harness notes short. If a note is reference material, save it as a doc in another folder; the model can still search for it.
08What we checked

Checked, and not claimed.

24 KB
Largest Harness note that still rides on every prompt
6/6
Harness filter tests pass
3
MCP dialects Luminair writes: SDK options, Codex overrides, Gemini settings
0
Connector secret values the MCP code puts on a command line: they travel by pipe, environment or variable reference

What this post does not claim

  • How many tokens your own connected servers cost. It depends on the servers, and our meter does not count schemas.
  • That the Anthropic figures apply to Luminair sessions. They are Anthropic's measurements of their own setups.
  • That every engine can use every connector. Codex skips custom-header servers, and engines without an MCP hook get none.

Sources

For more on what fills a model's context, read Your model didn't get dumber.

Connect once, use everywhere

Plug in a server and every engine that speaks MCP gets it, without its keys on a command line.

Download Luminair →