The short version
- Deep research agents, like Google's Gemini Deep Research, run for minutes in a data center and charge for the searches they make. Their limit is money and tool calls.
- Luminair's first multi-agent research ran on your Mac. Each agent was a full command-line engine process, about 330 MB before it did any work. Their limit was your memory.
- After a machine-lockup report in July 2026, an audit found the search agents could use every tool with no turn cap, because an allow-list does nothing when permissions are bypassed.
- Within about seven minutes, three commits made agents web-only and turn-capped, added a pre-flight memory gate and a live watchdog, and admitted agents one at a time. That guard now protects //deep-audit.
An agent that reads for you.
“Deep research” means an agent that takes one question, breaks it into many searches, reads what comes back, and writes a cited report. It can run for a long time, and every step is a decision.
On 11 December 2025, TechCrunch reported that Google had released a “reimagined” Gemini Deep Research that developers could embed in their own apps. The piece names the core risk of long agent runs, in which “many autonomous decisions are made over minutes, hours, or longer”: “The more choices an LLM has to make, the greater the chance that even one hallucinated choice will invalidate the entire output.”
In April 2026 Google followed with Deep Research and Deep Research Max. Max “leverages extended test-time compute to iteratively reason, search and refine the final report”, and Google pitches it for “asynchronous, background workflows such as a nightly cron job”.
These agents run on someone else's machines, so their cost shows up as a bill. Simon Willison tried OpenAI's o4-mini-deep-research through the API and added it up: 77 web search calls at about a cent each, plus tokens and code sessions, came to “$1.10” for one question. He also noted: “There's a limit on how many tool calls they can churn through in a single session.”
Five engines, one laptop.
Luminair's first version of this was a composer command, //research. A lead agent split your question into angles, several search agents each took one angle to the live web, and a final agent merged everything into one report with sources. A live card showed each agent working.
The difference from Google's version is where the agents ran. Each one was a separate Claude Code process started by the app, on your Mac, signed in to your account. The audit that followed measured each at about 330 MB of memory at baseline, before any real work. The search pool started five of them at once.
Five processes at 330 MB is over 1.6 GB before a single page is read. Add the pages, the growing contexts and whatever else the Mac was doing, and a smaller machine has nowhere to go but swap. In July 2026 came a report that a machine had locked up. The fixes landed on the morning of 21 July.
The allow-list that allowed everything.
The first commit after the report, 8ea3ebf7 at 08:13, is titled “Deep research: hard resource-safety pass after the machine-lockup report”. Its first finding was not about memory at all. The search agents ran with permissions bypassed, with every skill loaded, with every tool (including the shell and file writes), and with no limit on turns.
That was not the intent. The code passed a list of allowed tools. But the Agent SDK's permission docs are explicit that this combination does not restrict anything: “allowed_tools does not constrain bypassPermissions.” Their example: setting only Read as allowed alongside bypass mode “still approves every tool, including Bash, Write, and Edit.”
The fix switched research agents to dontAsk with a real whitelist. Search agents got only WebSearch and WebFetch and a cap of 16 turns. The lead and final agents got one turn each with every tool denied, since their job is only to plan and to write. No skills are loaded. The same commit cut the pool from five agents to three, allowed only one research run on the whole machine at a time (panes had been able to stack them), and made quitting the app or closing its window abort every running agent.
Other kinds of run in the app, like branch-out, kept their normal permissions. The lockdown applies to runs that are supposed to be narrow.
Measure, then admit.
Four minutes later, commit 9641791b added the rule in its own words: research “must never drown the Mac”. Three minutes after that, 4137c6ed added the rule that shaped everything since: “never start at full parallelism.”
1 · The pre-flight gate
Before any agent starts, the app reads the machine. If available memory is under 1,200 MB, or the one-minute load average is already over twice the number of cores, the run does not start and you see why: the Mac “only has” so much memory “available right now”, close some apps and try again.
“Available” needed care. Node's built-in free-memory number counts only pages that are completely unused, and on a healthy Mac that number is often small, since macOS keeps memory busy with caches it can give back. The commit calls it “crying-wolf freemem”. The app instead reads vm_stat and adds free, inactive, purgeable and speculative pages. If it cannot read the numbers, it reports unknown rather than raising a false alarm.
2 · The ramp
Agents are admitted one at a time. The first starts alone, since the gate already cleared the machine. Each next one waits ten seconds, re-measures, and starts only if at least 1,500 MB is available and load is under twice the cores, up to three. If the test fails, the ramp holds and tests again rather than failing the run. If the machine cannot be measured, no extra agent is added. The commit says the algorithm was checked against a scripted machine that went healthy, low on memory, overloaded and healthy again: it started at one, held twice, and resumed to three.
3 · The watchdog
While a run is active, a timer samples every three seconds: one ps call, from which the app builds its own process tree, and one vm_stat. It stops every agent at once if the agents together pass a line, if the whole Mac runs short, or if load passes three times the cores, and the card says “Emergency stop” with the reason. The timer exists only while a run does, so the app stays idle at idle.
The first lines were fixed: 3,500 MB for the agent tree and 600 MB of available memory. Four days later that was changed, because a flat 3.5 GB “killed research at 3.5GB on a big Mac where that's trivial”. Now the agent tree may use up to 80% of total RAM, never less than 3,500 MB, and the run stops if the whole Mac passes 90% used, with a 500 MB backstop.
The guard outlived the feature.
The names moved around quickly after that morning. That evening, in a change that also set up the Perplexity connector's own command, //research was renamed //deep-research. On 27 July our home-grown web research pipeline was retired in favour of Claude's own deep-research skill, and on 2 August the //deep-research shortcut came back to start that skill, with a live card that mirrors each search and source.
The guard stayed. The same day the pipeline was retired, a new command, //deep-audit, pointed the multi-agent machinery inward: a lead agent maps your codebase into surfaces, read-only agents inspect one surface each, and a final agent writes one ranked report with a fix plan. It uses the same pre-flight gate, ramp, watchdog and one-run-at-a-time lock. Its agents get Read, Grep, Glob and Bash in dontAsk mode, the final agent only the first three.
The ramp idea also spread. Branch-out, which runs several tasks at once, now uses its own adaptive pool that “mirrors the research ramp”: start one agent, let it settle, re-measure, and admit the next only with real headroom, with a faster start when the Mac clearly has room.
Watch the meter climb.
- 1In a session opened on your project folder, type //deep-audit, or //deep-audit followed by an area to focus on, and press Enter.
- 2A live card appears with the memory meter: RAM in use, the agents' share, load, and the auto-stop line. Agents join one at a time as the Mac allows.
- 3For research on the web, type //deep-research followed by your question. That runs Claude's own deep-research skill.
If the Mac is already short on memory, the audit will not start, and it tells you so. Only one heavy multi-agent run can be active at a time.
Checked, and not claimed.
What this post does not claim
- The details of the lockup itself: which Mac, how much memory, or what else was running. The commits record the report, not the machine.
- That the guard applies to Claude's own deep-research skill.
//deep-researchnow starts that skill; the pre-flight gate, ramp and watchdog described here protect//deep-audit. - That the thresholds are tuned for every Mac. They are round numbers chosen that morning and adjusted once. We found no dedicated unit test for them; the commits describe live checks on one machine and a scripted test of the ramp.
Sources
- TechCrunch · 11 December 2025Google launched its deepest AI research agent yet, on the same day OpenAI dropped GPT-5.2
- Google · 21 April 2026Deep Research Max: a step change for autonomous research agents
- Simon Willison's TILs · October 2025Exploring OpenAI's deep research API model o4-mini-deep-research
- Claude Code DocsAgent SDK: Configure permissions
For another run that could take a Mac down with it, read Many agents, one Mac.
Heavy runs, light touch
Run //deep-audit and watch agents join only as your Mac has room for them.