← Blog Systems & Memory 26 September 2026 8 min read

When every model fails at once, check the network first.

In October 2025 one empty DNS record in a single AWS region knocked apps offline all over the world, all at the same time. In September 2026 we lived through a tiny version of that on one Mac: every model broke in the same minutes, and every one of them blamed something different. None of them said the word offline. Now Luminair does.

Fig 01  Five messages, one cause17 September 2026 · schematic
WHAT EACH ONE SAID CLAUDE “No response from API” CODEX timed out on its model catalog ANTIGRAVITY streamed nothing for 60 seconds GOOGLE'S CLIENT “not logged in” LUMINAIR'S OWN RELAY POLL “fetch failed” THE ACTUAL CAUSE No internet One Mac, one connection, about 18 minutes WHAT LUMINAIR SAYS NOW NO INTERNET This Mac has had no internet since HH:MM. Codex could not reach its server. Reconnect and send again. Same wording for every CLI engine. Left: the symptoms as recorded in a code comment in desktop/main.js. Right: the message built in desktop/engines/runner-cli.js.
Wording on the left comes from the comment above the network-state code in desktop/main.js, written the day of the drop. Quoted phrases are the engines' own words; the other two are our summary of what they reported. The time on the right is filled in from the start of the outage.

The short version

  1. When many unrelated things break in the same minute, they usually share one cause further upstream. On 20 October 2025 that cause was an empty DNS record in AWS us-east-1.
  2. On 17 September 2026 our version was smaller: the Mac lost its internet for about 18 minutes, and Claude, Codex and Antigravity each failed with a different message. It read as every model being broken.
  3. Luminair already checks the cloud every 5 seconds for prompts sent from your phone. That poll is now also an internet heartbeat.
  4. When those checks are failing, a failed turn now says “This Mac has had no internet since HH:MM” instead of the engine's own confusing wording. Usage limits and sign-in problems are still reported as themselves.
02One record, many outages

The day the phonebook went blank.

Late on 19 October 2025, Pacific time, things started failing across the internet. ABC News reported that the AWS outage “disrupted hundreds of other global platforms, including Robinhood, Snapchat, Roblox and Perplexity.” Perplexity's chief executive, Aravind Srinivas, said he believed the root cause was an AWS issue. CNN wrote that “people couldn't order food, communicate with hospital networks, access mobile banking, or connect with their security systems and smart home devices.”

From the outside, that looked like dozens of separate companies having a bad day together. From the inside, it was one thing. AWS's own summary is specific: “The root cause of this issue was a latent race condition in the DynamoDB DNS management system that resulted in an incorrect empty DNS record.”

DNS is how a computer turns a name into an address. When the record for DynamoDB in Northern Virginia went empty, AWS says “all systems needing to connect to the DynamoDB service in the N. Virginia (us-east-1) Region via the public endpoint immediately began experiencing DNS failures and failed to connect.” The DynamoDB part lasted from 11:48 PM on 19 October to 2:40 AM on 20 October. But other AWS services that depend on it kept failing for much longer: the summary gives EC2 until 1:50 PM and Lambda until 2:15 PM that afternoon.

CNN's explainer quoted Angelique Medina of Cisco's ThousandEyes: “The analogy of a telephone book is pretty apt in that the folks on the other line are there, but if you don't know how to reach them, then you have a problem.”

When everything breaks in the same minute, the question is not what is wrong with each thing. It is what they all share.

That is the useful lesson for anyone who works with several AI tools at once. Each tool reports its own symptom. None of them can see the shared cause. The person looking at five red messages has to do the joining up.

03Our eighteen minutes

“No models are working right now.”

Luminair runs many AI engines side by side: Claude, Codex, Antigravity, Gemini and others, most of them through their own command-line tools. On 17 September 2026, the report that reached us was one line, now quoted in the code: “No models are working right now.”

It looked like a disaster in the harness. Every engine was failing, in the same minutes, and each one said something different. The comment we wrote into desktop/main.js that day lists them:

CLAUDE

“No response from API”

Reads like the provider is down.

CODEX

Timed out on its model catalog

Reads like a Codex bug.

AGY

Streamed nothing for 60 seconds

Antigravity hit Luminair's start watchdog. Reads like a hang.

GOOGLE

“Not logged in”

Reads like a broken sign-in. The sign-in was fine; refreshing its token needs the network.

Underneath, our own relay poll was logging fetch failed for the whole stretch, about 18 minutes. The comment ends: “Five messages, one cause, and nobody said "offline".”

Nothing was wrong with any model, any account or any line of Luminair's engine code. The Mac had simply lost its connection. But the messages pointed in four different directions, and the fifth, the one that actually named the problem, was only in a debug log.

04Why nobody says offline

Every tool reports its own symptom.

This is not carelessness on the part of each vendor. A command-line tool knows what it tried and what went wrong for it. It does not know whether the whole machine is cut off, whether a single server is down, or whether the problem is its own.

So each one describes the failure from where it stands. A tool waiting for a reply says there was no response. A tool that fetches a list first says the list timed out. A tool that refreshes a login first says you are not logged in, which is the most misleading of all, because it sends you off to fix an account that is working.

An app that wraps many engines inherits all of these voices. If it just passes them through, a single network drop looks like five separate faults. You could spend twenty minutes signing out and back in, restarting tools and reading status pages before you think to check the Wi-Fi.

The design questionLuminair cannot ask the engines whether the Mac is online. It has to know that itself, from a check that is already running, without adding any new traffic.
05A heartbeat we already had

The poll that was already running.

Luminair has a phone app. When you send a prompt from your phone, it lands in a cloud inbox, and the Mac picks it up. To make that feel instant, a signed-in Mac lists its inbox on a fixed 5 second timer. The code is blunt about why it never slows down: “their latency IS the phone experience”.

That means the app is already asking the internet a question every few seconds. The fix was to start listening to the answer. The inbox listing now records two things:

  • Any reply from the server, even an HTTP error, counts as a success. If a server answered, the Mac is online. The timestamp goes into _netOkAt.
  • A thrown network error counts as a failure, but only if it looks like one: fetch failed, ENOTFOUND, EAI_AGAIN, ETIMEDOUT, ENETUNREACH and a few more. The first failure after a success marks the start of a streak.
Fig 02  How the heartbeat decidesSchematic · one dot per 5 s check
streak starts any answer clears it macOfflineSince() returns the time of the first cross returns 0 STALENESS RULE no check at all (signed out, app asleep) after 3 minutes: returns 0 A verdict is never older than the last failed check plus 3 minutes.
The logic of noteNetFail and macOfflineSince in desktop/main.js. The spacing is illustrative; the 5 second timer, the error list and the 3 minute limit are from the code.

A small function, macOfflineSince(), turns that into one answer: the time the current streak of failed checks began, or zero. It returns zero the moment any check succeeds. It also returns zero if no check has failed in the last three minutes, so a Mac that was offline an hour ago and then stopped polling never gets told it is still offline. The comment says it plainly: “never a stale verdict.”

The whole thing is about a dozen lines. It adds no requests. It reuses traffic the app was already sending for a different reason.

06Naming the cause

Say the boring thing first.

The second half lives in the shared runner that drives Luminair's command-line engines: Codex, Antigravity, Gemini, Claude through its command-line tool, and most of the others. When a turn fails for good, the runner builds the message you see. It now asks the app one question first: is this Mac offline right now?

Fig 03  Which message you getFrom runner-cli.js · simplified
A turn fails Automatic recovery first retries, fresh thread, spare account Still failing: build the final message Limit, blocked or sign-in? yes Reported as itself those are not the network no Checks failing right now? yes “This Mac has had no internet since HH:MM.” shown as a notice card, not an error no The engine's own words
Order of checks in the final error path of desktop/engines/runner-cli.js. The recovery steps listed in the top middle box depend on the kind of failure; not every turn gets every one.

The order matters. Every automatic recovery the runner already had still runs first. The offline wording only replaces the last message, the one you would otherwise read. And three kinds of failure are deliberately left alone: a usage limit, a blocked account and a failed sign-in. The code comment gives the reason: “a limit or a broken sign-in is still reported as itself because those are not the network.” If your plan ran out while the Wi-Fi also dropped, you should hear about the plan.

For everything else, the message is built from the streak's start time and the engine's name:

desktop/engines/runner-cli.jsfinal error path
if (offSince && failureKind !== 'account-limit' && failureKind !== 'account-blocked' && failureKind !== 'auth') {
  offline = true;
  // hh = the streak's start time as HH:MM
  msg = 'This Mac has had no internet since ' + hh + '. ' + spec.label + ' could not reach its server. Reconnect and send again.';
}

The error also carries an offline flag. On 26 September we used it to change how the message looks. An internet drop is not a fault in anything you can fix inside the app, so it no longer appears as an error banner. It appears as a calm notice card with the heading No internet. The comment next to that change quotes the reason: “This looks SCARY, it's just internet”.

Luminair · a session while offline
Run the test suite and fix whatever fails.
No internet×
This Mac has had no internet since 09:41. Codex could not reach its server. Reconnect and send again.
Drawn from the notice card in desktop/renderer.js: a dot, the heading “No internet”, a dismiss button and the runner's message. The prompt and the time are illustrative. Unlike passing notices, it is not cleared automatically; it stays until you dismiss it.
07When it happens to you

A two minute triage.

The heartbeat only covers what Luminair can see. The general habit is worth more than the feature, and it works in any tool.

  1. 1Count the failures. If two or more different models failed in the same few minutes, suspect something they share before suspecting each of them.
  2. 2Check the closest shared thing first: your own connection. Open any website. In Luminair, look for the No internet card in the session.
  3. 3If your connection is fine, check the next shared thing out: the providers' status pages, and whether a big cloud region is having a day like 20 October 2025.
  4. 4Only then treat each failure on its own: sign-ins, limits, a specific model. And be suspicious of “not logged in” during an outage. It may just mean the tool could not refresh its token.
08What we checked

Checked, and not claimed.

2/2
Runner tests pass: a failure while offline names the outage; a usage limit while offline is still the limit
5 s
Interval of the inbox poll that doubles as the heartbeat, from the timer in desktop/main.js
3 min
Longest a verdict can outlive the last failed check
0
Extra network requests added for the heartbeat

What this post does not claim

  • That Luminair detects every kind of outage. It detects one: this Mac cannot reach the internet. A provider that is down while your connection is fine still shows the provider's own message.
  • That every engine is covered. The rule lives in the shared command-line runner. Claude turns that run through the Agent SDK lane do not use it.
  • That it works while you are signed out. The heartbeat is the inbox poll, and the poll only runs for a Mac signed in to a Luminair account.
  • Anything about the AWS outage beyond what AWS, CNN and ABC News published. We link to them below.

For a related story about the harness looking like the model, read Your model didn't get dumber. Your harness might have.

Sources

Fewer red herrings

Run several models side by side, and let the app tell you when the problem is the connection.

Download Luminair →