← Blog Systems & Memory 2 August 2026 9 min read

Deep research that can't drown the Mac

Cloud research agents cost money for every search they run. Run the same idea on a laptop and the bill arrives as memory. In July 2026 Luminair's own multi-agent research locked up a Mac. What followed was a morning of fixes that now guard every heavy run in the app.

Fig 01  Three layers around one runSchematic · thresholds from RESEARCH_GUARD in desktop/main.js
LAYER 1 · BEFORE Pre-flight gate Refuse to start if available RAM is under 1,200 MB or load is already over 2 × the core count. Zero agents spawned. LAYER 2 · RAMP Earn each agent 1 wait 10 s, test 12 wait 10 s, test hold, re-test 123 cap Test: ≥ 1,500 MB available and load < 2 × cores. Cannot measure: add nothing. One cited report or one ranked audit, from a final agent with its tools limited to the job LAYER 3 · THE WHOLE TIME, EVERY 3 SECONDS Watchdog one ps and one vm_stat per sample; the timer exists only while a run does Agent treeover 80% of RAM(never below 3,500 MB) Whole Macover 90% used(500 MB backstop) Loadover 3 × cores stop all
Numbers are the current constants in desktop/main.js. The watchdog's lines were fixed megabytes on 21 July 2026 and became shares of total RAM on 25 July, so a large Mac is not stopped at 3.5 GB. Agent counts are drawn from the ramp's rules; a real run may hold at one or two.

The short version

  1. Deep research agents, like Google's Gemini Deep Research, run for minutes in a data center and charge for the searches they make. Their limit is money and tool calls.
  2. Luminair's first multi-agent research ran on your Mac. Each agent was a full command-line engine process, about 330 MB before it did any work. Their limit was your memory.
  3. After a machine-lockup report in July 2026, an audit found the search agents could use every tool with no turn cap, because an allow-list does nothing when permissions are bypassed.
  4. Within about seven minutes, three commits made agents web-only and turn-capped, added a pre-flight memory gate and a live watchdog, and admitted agents one at a time. That guard now protects //deep-audit.
02Deep research, in the cloud

An agent that reads for you.

“Deep research” means an agent that takes one question, breaks it into many searches, reads what comes back, and writes a cited report. It can run for a long time, and every step is a decision.

On 11 December 2025, TechCrunch reported that Google had released a “reimagined” Gemini Deep Research that developers could embed in their own apps. The piece names the core risk of long agent runs, in which “many autonomous decisions are made over minutes, hours, or longer”: “The more choices an LLM has to make, the greater the chance that even one hallucinated choice will invalidate the entire output.”

In April 2026 Google followed with Deep Research and Deep Research Max. Max “leverages extended test-time compute to iteratively reason, search and refine the final report”, and Google pitches it for “asynchronous, background workflows such as a nightly cron job”.

These agents run on someone else's machines, so their cost shows up as a bill. Simon Willison tried OpenAI's o4-mini-deep-research through the API and added it up: 77 web search calls at about a cent each, plus tokens and code sessions, came to “$1.10” for one question. He also noted: “There's a limit on how many tool calls they can churn through in a single session.”

In a data center, a runaway agent costs money. On a laptop, it costs the laptop.
03Ours ran on your Mac

Five engines, one laptop.

Luminair's first version of this was a composer command, //research. A lead agent split your question into angles, several search agents each took one angle to the live web, and a final agent merged everything into one report with sources. A live card showed each agent working.

The difference from Google's version is where the agents ran. Each one was a separate Claude Code process started by the app, on your Mac, signed in to your account. The audit that followed measured each at about 330 MB of memory at baseline, before any real work. The search pool started five of them at once.

Five processes at 330 MB is over 1.6 GB before a single page is read. Add the pages, the growing contexts and whatever else the Mac was doing, and a smaller machine has nowhere to go but swap. In July 2026 came a report that a machine had locked up. The fixes landed on the morning of 21 July.

04The lockup and the audit

The allow-list that allowed everything.

The first commit after the report, 8ea3ebf7 at 08:13, is titled “Deep research: hard resource-safety pass after the machine-lockup report”. Its first finding was not about memory at all. The search agents ran with permissions bypassed, with every skill loaded, with every tool (including the shell and file writes), and with no limit on turns.

That was not the intent. The code passed a list of allowed tools. But the Agent SDK's permission docs are explicit that this combination does not restrict anything: “allowed_tools does not constrain bypassPermissions.” Their example: setting only Read as allowed alongside bypass mode “still approves every tool, including Bash, Write, and Edit.”

Fig 02  Where a tool call goesSchematic · from the SDK permission docs
BEFORE · ALLOW-LIST + bypassPermissions Bash call on the allow-list? no mode: bypass runs SINCE 21 JUL · ALLOW-LIST + dontAsk Bash call on the allow-list? no mode: dontAsk denied In both rows, WebSearch and WebFetch are on the list and approved directly. Only the fall-through changes.
The allow-list is the same in both rows. What changed is the permission mode that decides everything not on the list. The docs recommend exactly this pairing “For a locked-down agent”.

The fix switched research agents to dontAsk with a real whitelist. Search agents got only WebSearch and WebFetch and a cap of 16 turns. The lead and final agents got one turn each with every tool denied, since their job is only to plan and to write. No skills are loaded. The same commit cut the pool from five agents to three, allowed only one research run on the whole machine at a time (panes had been able to stack them), and made quitting the app or closing its window abort every running agent.

Other kinds of run in the app, like branch-out, kept their normal permissions. The lockdown applies to runs that are supposed to be narrow.

05Three layers

Measure, then admit.

Four minutes later, commit 9641791b added the rule in its own words: research “must never drown the Mac”. Three minutes after that, 4137c6ed added the rule that shaped everything since: “never start at full parallelism.”

1 · The pre-flight gate

Before any agent starts, the app reads the machine. If available memory is under 1,200 MB, or the one-minute load average is already over twice the number of cores, the run does not start and you see why: the Mac “only has” so much memory “available right now”, close some apps and try again.

“Available” needed care. Node's built-in free-memory number counts only pages that are completely unused, and on a healthy Mac that number is often small, since macOS keeps memory busy with caches it can give back. The commit calls it “crying-wolf freemem”. The app instead reads vm_stat and adds free, inactive, purgeable and speculative pages. If it cannot read the numbers, it reports unknown rather than raising a false alarm.

2 · The ramp

Agents are admitted one at a time. The first starts alone, since the gate already cleared the machine. Each next one waits ten seconds, re-measures, and starts only if at least 1,500 MB is available and load is under twice the cores, up to three. If the test fails, the ramp holds and tests again rather than failing the run. If the machine cannot be measured, no extra agent is added. The commit says the algorithm was checked against a scripted machine that went healthy, low on memory, overloaded and healthy again: it started at one, held twice, and resumed to three.

3 · The watchdog

While a run is active, a timer samples every three seconds: one ps call, from which the app builds its own process tree, and one vm_stat. It stops every agent at once if the agents together pass a line, if the whole Mac runs short, or if load passes three times the cores, and the card says “Emergency stop” with the reason. The timer exists only while a run does, so the app stays idle at idle.

The first lines were fixed: 3,500 MB for the agent tree and 600 MB of available memory. Four days later that was changed, because a flat 3.5 GB “killed research at 3.5GB on a big Mac where that's trivial”. Now the agent tree may use up to 80% of total RAM, never less than 3,500 MB, and the run stops if the whole Mac passes 90% used, with a 500 MB backstop.

//deep-audit · live card
Deep audit
% RAM used · agents % () · load · auto kill 90%
Redrawn from memMeterHtml in desktop/renderer.js, added on 25 July. Numbers are left blank on purpose. The hatched bar is whole-Mac memory in use; the vertical line is the stop line. In the app the bar also changes tone as it nears the line; here that is left out. The rows stand for agents: filled dot, admitted; hollow, not yet admitted. Their labels are left blank.
06Where it lives now

The guard outlived the feature.

The names moved around quickly after that morning. That evening, in a change that also set up the Perplexity connector's own command, //research was renamed //deep-research. On 27 July our home-grown web research pipeline was retired in favour of Claude's own deep-research skill, and on 2 August the //deep-research shortcut came back to start that skill, with a live card that mirrors each search and source.

The guard stayed. The same day the pipeline was retired, a new command, //deep-audit, pointed the multi-agent machinery inward: a lead agent maps your codebase into surfaces, read-only agents inspect one surface each, and a final agent writes one ranked report with a fix plan. It uses the same pre-flight gate, ramp, watchdog and one-run-at-a-time lock. Its agents get Read, Grep, Glob and Bash in dontAsk mode, the final agent only the first three.

The ramp idea also spread. Branch-out, which runs several tasks at once, now uses its own adaptive pool that “mirrors the research ramp”: start one agent, let it settle, re-measure, and admit the next only with real headroom, with a faster start when the Mac clearly has room.

The lesson we tookParallelism is not a setting; it is something the machine earns, one agent at a time. And an allow-list is only as strong as the mode it sits in. Read the permission docs for what happens to everything not on the list.
07Find it in the app

Watch the meter climb.

  1. 1In a session opened on your project folder, type //deep-audit, or //deep-audit followed by an area to focus on, and press Enter.
  2. 2A live card appears with the memory meter: RAM in use, the agents' share, load, and the auto-stop line. Agents join one at a time as the Mac allows.
  3. 3For research on the web, type //deep-research followed by your question. That runs Claude's own deep-research skill.

If the Mac is already short on memory, the audit will not start, and it tells you so. Only one heavy multi-agent run can be active at a time.

08What we checked

Checked, and not claimed.

~330 MB
Baseline memory per agent process, as measured in the 21 July audit
5 → 3
Maximum parallel agents, now reached only through the ramp
3 s
Watchdog sampling interval, only while a run is active
~7 min
From the first safety commit (08:13) to the ramp (08:20), 21 July 2026

What this post does not claim

  • The details of the lockup itself: which Mac, how much memory, or what else was running. The commits record the report, not the machine.
  • That the guard applies to Claude's own deep-research skill. //deep-research now starts that skill; the pre-flight gate, ramp and watchdog described here protect //deep-audit.
  • That the thresholds are tuned for every Mac. They are round numbers chosen that morning and adjusted once. We found no dedicated unit test for them; the commits describe live checks on one machine and a scripted test of the ramp.

Sources

For another run that could take a Mac down with it, read Many agents, one Mac.

Heavy runs, light touch

Run //deep-audit and watch agents join only as your Mac has room for them.

Download Luminair →