The short version
- When many unrelated things break in the same minute, they usually share one cause further upstream. On 20 October 2025 that cause was an empty DNS record in AWS us-east-1.
- On 17 September 2026 our version was smaller: the Mac lost its internet for about 18 minutes, and Claude, Codex and Antigravity each failed with a different message. It read as every model being broken.
- Luminair already checks the cloud every 5 seconds for prompts sent from your phone. That poll is now also an internet heartbeat.
- When those checks are failing, a failed turn now says “This Mac has had no internet since HH:MM” instead of the engine's own confusing wording. Usage limits and sign-in problems are still reported as themselves.
The day the phonebook went blank.
Late on 19 October 2025, Pacific time, things started failing across the internet. ABC News reported that the AWS outage “disrupted hundreds of other global platforms, including Robinhood, Snapchat, Roblox and Perplexity.” Perplexity's chief executive, Aravind Srinivas, said he believed the root cause was an AWS issue. CNN wrote that “people couldn't order food, communicate with hospital networks, access mobile banking, or connect with their security systems and smart home devices.”
From the outside, that looked like dozens of separate companies having a bad day together. From the inside, it was one thing. AWS's own summary is specific: “The root cause of this issue was a latent race condition in the DynamoDB DNS management system that resulted in an incorrect empty DNS record.”
DNS is how a computer turns a name into an address. When the record for DynamoDB in Northern Virginia went empty, AWS says “all systems needing to connect to the DynamoDB service in the N. Virginia (us-east-1) Region via the public endpoint immediately began experiencing DNS failures and failed to connect.” The DynamoDB part lasted from 11:48 PM on 19 October to 2:40 AM on 20 October. But other AWS services that depend on it kept failing for much longer: the summary gives EC2 until 1:50 PM and Lambda until 2:15 PM that afternoon.
CNN's explainer quoted Angelique Medina of Cisco's ThousandEyes: “The analogy of a telephone book is pretty apt in that the folks on the other line are there, but if you don't know how to reach them, then you have a problem.”
That is the useful lesson for anyone who works with several AI tools at once. Each tool reports its own symptom. None of them can see the shared cause. The person looking at five red messages has to do the joining up.
“No models are working right now.”
Luminair runs many AI engines side by side: Claude, Codex, Antigravity, Gemini and others, most of them through their own command-line tools. On 17 September 2026, the report that reached us was one line, now quoted in the code: “No models are working right now.”
It looked like a disaster in the harness. Every engine was failing, in the same minutes, and each one said something different. The comment we wrote into desktop/main.js that day lists them:
“No response from API”
Reads like the provider is down.
Timed out on its model catalog
Reads like a Codex bug.
Streamed nothing for 60 seconds
Antigravity hit Luminair's start watchdog. Reads like a hang.
“Not logged in”
Reads like a broken sign-in. The sign-in was fine; refreshing its token needs the network.
Underneath, our own relay poll was logging fetch failed for the whole stretch, about 18 minutes. The comment ends: “Five messages, one cause, and nobody said "offline".”
Nothing was wrong with any model, any account or any line of Luminair's engine code. The Mac had simply lost its connection. But the messages pointed in four different directions, and the fifth, the one that actually named the problem, was only in a debug log.
Every tool reports its own symptom.
This is not carelessness on the part of each vendor. A command-line tool knows what it tried and what went wrong for it. It does not know whether the whole machine is cut off, whether a single server is down, or whether the problem is its own.
So each one describes the failure from where it stands. A tool waiting for a reply says there was no response. A tool that fetches a list first says the list timed out. A tool that refreshes a login first says you are not logged in, which is the most misleading of all, because it sends you off to fix an account that is working.
An app that wraps many engines inherits all of these voices. If it just passes them through, a single network drop looks like five separate faults. You could spend twenty minutes signing out and back in, restarting tools and reading status pages before you think to check the Wi-Fi.
The poll that was already running.
Luminair has a phone app. When you send a prompt from your phone, it lands in a cloud inbox, and the Mac picks it up. To make that feel instant, a signed-in Mac lists its inbox on a fixed 5 second timer. The code is blunt about why it never slows down: “their latency IS the phone experience”.
That means the app is already asking the internet a question every few seconds. The fix was to start listening to the answer. The inbox listing now records two things:
- Any reply from the server, even an HTTP error, counts as a success. If a server answered, the Mac is online. The timestamp goes into
_netOkAt. - A thrown network error counts as a failure, but only if it looks like one:
fetch failed,ENOTFOUND,EAI_AGAIN,ETIMEDOUT,ENETUNREACHand a few more. The first failure after a success marks the start of a streak.
noteNetFail and macOfflineSince in desktop/main.js. The spacing is illustrative; the 5 second timer, the error list and the 3 minute limit are from the code.A small function, macOfflineSince(), turns that into one answer: the time the current streak of failed checks began, or zero. It returns zero the moment any check succeeds. It also returns zero if no check has failed in the last three minutes, so a Mac that was offline an hour ago and then stopped polling never gets told it is still offline. The comment says it plainly: “never a stale verdict.”
The whole thing is about a dozen lines. It adds no requests. It reuses traffic the app was already sending for a different reason.
Say the boring thing first.
The second half lives in the shared runner that drives Luminair's command-line engines: Codex, Antigravity, Gemini, Claude through its command-line tool, and most of the others. When a turn fails for good, the runner builds the message you see. It now asks the app one question first: is this Mac offline right now?
The order matters. Every automatic recovery the runner already had still runs first. The offline wording only replaces the last message, the one you would otherwise read. And three kinds of failure are deliberately left alone: a usage limit, a blocked account and a failed sign-in. The code comment gives the reason: “a limit or a broken sign-in is still reported as itself because those are not the network.” If your plan ran out while the Wi-Fi also dropped, you should hear about the plan.
For everything else, the message is built from the streak's start time and the engine's name:
if (offSince && failureKind !== 'account-limit' && failureKind !== 'account-blocked' && failureKind !== 'auth') { offline = true; // hh = the streak's start time as HH:MM msg = 'This Mac has had no internet since ' + hh + '. ' + spec.label + ' could not reach its server. Reconnect and send again.'; }
The error also carries an offline flag. On 26 September we used it to change how the message looks. An internet drop is not a fault in anything you can fix inside the app, so it no longer appears as an error banner. It appears as a calm notice card with the heading No internet. The comment next to that change quotes the reason: “This looks SCARY, it's just internet”.
A two minute triage.
The heartbeat only covers what Luminair can see. The general habit is worth more than the feature, and it works in any tool.
- 1Count the failures. If two or more different models failed in the same few minutes, suspect something they share before suspecting each of them.
- 2Check the closest shared thing first: your own connection. Open any website. In Luminair, look for the No internet card in the session.
- 3If your connection is fine, check the next shared thing out: the providers' status pages, and whether a big cloud region is having a day like 20 October 2025.
- 4Only then treat each failure on its own: sign-ins, limits, a specific model. And be suspicious of “not logged in” during an outage. It may just mean the tool could not refresh its token.
Checked, and not claimed.
What this post does not claim
- That Luminair detects every kind of outage. It detects one: this Mac cannot reach the internet. A provider that is down while your connection is fine still shows the provider's own message.
- That every engine is covered. The rule lives in the shared command-line runner. Claude turns that run through the Agent SDK lane do not use it.
- That it works while you are signed out. The heartbeat is the inbox poll, and the poll only runs for a Mac signed in to a Luminair account.
- Anything about the AWS outage beyond what AWS, CNN and ABC News published. We link to them below.
For a related story about the harness looking like the model, read Your model didn't get dumber. Your harness might have.
Sources
- Amazon Web Services · October 2025Summary of the Amazon DynamoDB Service Disruption in the Northern Virginia (US-EAST-1) Region
- CNN Business · 25 October 2025How a tiny bug spiraled into a massive outage that took down the internet
- ABC News · 20 October 2025Amazon cloud unit AWS experiences major global outage disrupting hundreds of platforms
- Hacker News · 20 October 2025AWS multiple services outage in us-east-1
Fewer red herrings
Run several models side by side, and let the app tell you when the problem is the connection.