← Blog Security & Trust 10 September 2026 9 min read

The test that killed every process on the Mac

In late 2025 two coding agents deleted far more than they were asked to. On 7 September 2026 our own test suite did something similar to the Mac it ran on, five times in one day. The cause was one number. Here is the bug, the rule that replaced it, and what Luminair puts between an agent and your files.

Fig 01  What kill() does with the number you pass itSchematic · macOS kill(2)
THE CALL WHO GETS THE SIGNAL APP CHROME WAKE RUNNERS TEST GATE kill(4821) One process: the one whose id is 4821. The gate's shell, and nothing else. kill(-4821) Process group 4821: the shell and its children. What a cleanup means to do. Nothing a gate started survives. kill(0) The sender's own process group. A pid of 0, negated, is still 0: the test run signals itself. kill(-1) Every process with your user id, except the sender. What a fake pid of 1 became on 7 September. SPARED
Schematic. The four rules are from the macOS kill(2) manual page; the processes are the ones the fix commit names for 7 September. Filled dot: receives the signal. Hollow dot: untouched. The pid 4821 is made up.

The short version

  1. Agents and scripts run with your permissions. When they get a target wrong, the operating system does not ask whether you meant it.
  2. In our case a test gave a cleanup routine a fake process id of 1. Negated, that became kill(-1): signal every process you own. Each run of the core suite closed the app, Chrome and the build runners.
  3. Every place in Luminair that kills a process group now requires an integer pid above 1, and a regression test proves 1, 0 and a missing pid never reach the kernel as a group kill.
  4. For what agents do to files, Luminair snapshots the project folder before every turn (//rewind) and offers an OS-level fence (/sandbox) that shows a card when it blocks a write.
02Two cleanups gone wrong

“Clear the cache” is a dangerous sentence.

At the start of December 2025, The Register reported that Google's Antigravity agent had wiped a user's entire D: drive. The user had asked it to clear a project cache. According to Tom's Hardware, the agent later explained that its rmdir command “appears to have incorrectly targeted the root of your D: drive instead of the specific project folder”, and that because it used the quiet flag, the files skipped the Recycle Bin. It ended with: “I am deeply, deeply sorry. This is a critical failure on my part.” The user had been running in Turbo mode, which, in The Register's words, “lets the Antigravity agent execute commands without user input”.

A week later came the case Docker's engineering blog calls the rm -rf ~/ incident. A developer asked Claude Code to clean up an old repository. The command it ran listed three project folders and then ~/, which the shell expands to the home directory. Docker's summary is blunt: “It was the AI coding agent doing exactly what it was told, in a way the user did not anticipate, with no architectural boundary to catch the mistake.”

Both stories share a shape. A routine request. A command that is valid syntax. One argument that points somewhere much bigger than intended. And a system that carries it out at full speed with the user's own rights.

The shell does not know which folder you meant. It only knows which one you named.
03Our own kill(-1)

No agent needed. Just a test.

It would be comfortable to treat these as model failures. Our own worst day with a destructive command had no model in it at all.

Luminair has goals with gates. A gate is a command you name, such as npm test, whose exit code decides whether a goal is done. The model's words never mark success. The gate runs in its own process group, so that when it finishes or times out, Luminair can kill the whole group and leave nothing running behind it. On Unix you kill a group by passing its id as a negative number: process.kill(-child.pid, 'SIGKILL').

The unit test for that code does not start a real shell. It hands in a stand-in child process, a “test double”, with a made-up pid. The made-up pid, added on 6 September, was 1. On its own that was harmless, because only the timeout path sent a group kill and the fake gate finished instantly. Then an in-progress change to the gate added a cleanup that kills the group every time a gate finishes, so no leftover process could survive a verification command. From then on, every run of that test called process.kill(-1, 'SIGKILL').

The manual page is precise about what that means. If you are not the superuser, “the signal is sent to all processes with the same uid as the user, excluding the process sending the signal.” Everything you are logged in as. The test runner survived its own call; almost nothing else did.

The fix commit on 7 September lists the damage: the app itself, Chrome, a wake agent, a local runner and the GitHub Actions runner, five times that day, and once a shutdown that stalled into a reboot.

Why it hid so wellSIGKILL cannot be caught or logged by the process that receives it. Every app simply vanished, with no error dialog and nothing in its own log. The only process that knew what happened was the one that sent the signal, and it was a test that passed.
04One rule, five places

A pid must earn a group kill.

Fixing the test double alone would have left the trap in place for the next one. So the commit went looking for every place that signals a process group, and found five. Each now checks the pid first and falls back to killing only the child process when the check fails.

Fig 02  Every group kill in the desktop appFrom commit 2db9c9c0
WhereWhat it stopsGuard now
lib/goal-gate.jsA gate command and anything it spawnedinteger, > 1
lib/cloud/workspace-tools.jsTool processes run for cloud workspacesinteger, > 1
lib/router.js · killTreeA routed engine and its childreninteger, > 1
engines/runner-cli.js · killTreeA command-line engine and its tool serversinteger, > 1
main.js · keepJobStopA background job, from the pid it saved> 1
Real code paths. If the check fails, the code signals only the child it holds, never a group. The Keep job reads its pid from a file, so a missing or zero value parses to 0 and is refused the same way.

The gate's version is the clearest:

desktop/lib/goal-gate.jsgroup kill guard
// A process group may be signalled as -pid only for a genuine child pid. pid 0 and 1
// (and anything non-numeric) turn kill(-pid) into a session-wide or group-wide kill.
function groupKillable(child) { const p = child && child.pid; return Number.isInteger(p) && p > 1; }

Two more changes came with it. The test double now uses a pid that can never belong to a real process, with a comment saying why. And a new test swaps out process.kill, runs the gate with fake pids of 1, 0 and undefined, and fails if any call ever arrives as -1, 0 or NaN. Guarding against a number is cheap. Remembering, a year from now, why the guard is there is the hard part, and the test is how that memory is kept.

05Rewind: undo for files

Assume the command will go wrong.

A pid check protects processes. For files, the honest assumption is that some command will eventually delete the wrong thing, and approval prompts will not always catch it. Luminair's answer, shipped with Harness v1 on 29 August 2026, is to make loss recoverable.

Before every turn, whichever engine runs it, Luminair takes a snapshot of the session's project folder into a separate git repository stored outside the project. Your own repo and its git status are not touched. The snapshot catches what shell commands change too, not only the model's file edits. Heavy build folders such as node_modules and dist are left out.

Fig 03  A snapshot before every turnSchematic · time →
SNAPSNAPSNAPPRE-RESTORE turn 1 · edits turn 2 · edits turn 3 · rm -rf files gone Restore state saved first, so the undo is undoable PROJECT FOLDER · SHADOW REPO OUTSIDE THE PROJECT
Schematic of the rules in desktop/lib/reversible.js. Filled squares are pre-turn snapshots; the hollow one is the state saved just before a restore. Turns less than 15 seconds apart share one snapshot, and a turn never waits more than 8 seconds for one.

Type //rewind and a sheet lists those snapshots, newest first, each labelled with the prompt that followed it. What changed since shows a file-by-file summary. Restore needs a second tap, then commits the current state before rolling the whole folder back, so a restore can itself be rewound.

It is important to say what Rewind does not cover, because the file says so too. It only protects the session's project folder. It refuses to snapshot your home folder, the disk root, containers like Documents and Desktop, and anything outside your home folder, such as an external drive. A command that reaches past the project, like a trailing ~/ or a wrong drive letter, reaches past Rewind too. Files created and deleted inside a single turn are also invisible to it.

06A fence that asks

Blocked, and told about it.

For the reach problem, Luminair has Protected Mode, opened with /sandbox. When it is on, a run still goes full speed, but the operating system stops it from writing outside the session's folder and the folders you allow. It is off by default since 10 September 2026: you switch it on per session, or a Team policy can require it.

A fence that fails silently has its own problem: the model hits the wall, cannot tell why, and starts improvising around it. So when a tool result comes back with an error like “operation not permitted” or “read-only file system” inside a fenced Claude run, Luminair shows one card for that run:

Luminair · fence-hit card
Protected Mode

The fence blocked a write outside this session's folder: /Users/you/other-project. Allow it and tell the model to try again, or leave it blocked.

Allow onceAlways for this sessionDeny
Drawn from the card's text in desktop/renderer.js; the path is an example. The grant covers the folder, not a single file, and takes effect from the next turn, because the blocked command has already failed. The buttons lock after one choice.

Not every engine can enforce the fence. The built-in Claude engine applies it through its SDK, and among the command-line engines only Codex enforces it, on macOS and Linux. If a session requires protection and its engine cannot enforce it, Luminair refuses to run the turn rather than quietly running it open. If protection is on but not required, the run goes ahead and the session says it is unfenced.

07Find it in the app

Two commands worth knowing.

  1. 1Pick a project folder on the left. After your next turn, type //rewind in the composer to see the snapshot timeline for that folder.
  2. 2On any snapshot, click What changed since to see the damage before you decide. Click Restore twice to roll the folder back.
  3. 3Type /sandbox to open the fence for the current session and switch it on. Add any other folders the session may write to.
08What we checked

Checked, and not claimed.

5
Group-kill call sites guarded in one commit, on the day of the incident
9/9
Goal-gate tests pass, including the pid 1, 0 and missing-pid regression
8 s
Longest a turn waits for its pre-turn snapshot before going ahead
3
Choices on a fence-hit card: once, this session, or deny

What this post does not claim

  • That Rewind or the fence would have stopped either public incident. Both reached beyond a project folder, which is exactly where Rewind stops.
  • That the kill(-1) bug ever reached a user. The fake pid and the new cleanup only met in the working tree of our build Mac, and the guard landed in the same commit that first recorded that cleanup.
  • That Protected Mode is on for you. It is opt-in per session unless a Team policy requires it.

Sources

Give your agents an undo button

Every turn is snapshotted first. Type //rewind to see the timeline.

Download Luminair →