The short version
- Luminair keeps your Macs and phones in sync through a small rooms worker on Cloudflare. Each device holds a live socket to it.
- On 23 September a merge-resolution commit silently replaced the worker's source with an older version that could not read the way Macs send their ticket.
- On 26 September a routine deploy shipped it. For about four hours, every Mac socket was refused. Phones kept working, and the app logged nothing.
- Now the only supported deploy runs a gate: a source check, the tests, a live handshake probe, and an automatic rollback if the probe fails. The app also falls back to the other handshake by itself.
Two outages with the same shape.
On 18 November 2025, Cloudflare's network began failing at 11:20 UTC. A permissions change in a database made a Bot Management “feature file” double in size, and “the larger-than-expected feature file was then propagated to all the machines that make up our network.” Cloudflare called it its worst outage since 2019. Among the follow-ups, it promised to harden “ingestion of Cloudflare-generated configuration files in the same way we would for user-generated input”.
Two and a half weeks later, on 5 December, it happened again, for about 25 minutes and roughly 28% of Cloudflare's HTTP traffic. The trigger this time was a change made through the global configuration system, which, in Cloudflare's words, “does not perform gradual rollouts, but rather propagates changes within seconds to the entire fleet”. The post draws the lesson itself: “In both cases, a deployment to help mitigate a security issue for our customers propagated to our entire network.”
The November post-mortem was one of the most read engineering write-ups of the year, with more than 1,400 points on Hacker News. The reason it landed is that the pattern is universal. Every team has some path where one change goes straight to everyone.
We are nowhere near Cloudflare's scale. But in September 2026 we found one of those paths in our own setup, and it ran on their platform.
A switchboard for your devices.
Luminair's desktop app runs the sessions. Your phone, your other Macs and your teammates watch and steer them. The live part of that runs through one Cloudflare Worker with two kinds of Durable Object: an account room, one per user, that every Mac and phone on the account joins, and a session room for each shared session. Prompts, replies, presence and queue changes travel as frames through those rooms.
To join a room, a device shows a ticket: a short, signed pass minted by our backend after it checks who you are. The worker verifies the signature and lets the socket in.
Where that ticket travels matters. Phones put it in the URL, as ?t=…. On 15 September, a security assessment of the computer-to-computer relay flagged exactly that as finding F-09: a bearer ticket in the URL query string, with request logging switched on. So the desktop app moved its ticket into the WebSocket subprotocol header instead. It offers two protocols, ow-ticket and ow-t.<ticket>, and a hardened worker reads the ticket from there and answers with ow-ticket.
url.searchParams.get('t') and nothing else, which is why phones, which use ?t=, were unaffected.A merge that looked routine.
In mid-September two lines of work touched the worker at the same time. One was the hardened relay from the 15 September assessment: ticket in the subprotocol, per-device message budgets, computer-to-computer transfer frames, structured refusal logs. The other was an older “relay phase one” lineage, which had its own feature, computers following a session.
They were never fully reconciled. Each version fails the other's tests: the hardened source fails four “following” tests in the account-feeds file, and the older one fails the hardening tests. Production had been running the hardened one.
On 23 September, a commit titled “Resolve merged relay and Scrapbook branches” settled a conflict in desktop/rooms/src/index.js by taking the older side. The diff removed ticketFromProtocols, the subprotocol echo, the rate budget and the transfer frames. Nobody deployed the worker that day, so nothing broke. The swapped source simply sat in the repository, looking like any other file.
Three days later, someone ran wrangler deploy from that folder for an unrelated change. Wrangler did exactly what it was told. It uploaded the source it found to every Cloudflare location at once. That is the whole point of Workers, and on this day it was the problem.
Nothing crashed. Nothing logged.
From 14:51 UTC to 18:54 UTC on 26 September, every Mac's account-room socket and every session-room socket failed its handshake. Three things made it hard to see.
The failure was before “open”
The shared socket code logged when a socket opened and when an open socket closed. A socket refused during the handshake never opened, so it produced neither line. It just retried, with backoff, forever.
Phones still worked
Phones send the ticket in the URL, which the older worker understood. Their own sockets stayed up the whole time.
A fallback hid the rest
With the live path down, phones still caught up through the slower Firestore polling path. So the symptom was not “sync is broken”. It was “the phone syncs, but minutes late”.
It was found by chasing that lag. The fix was to restore the hardened source from the commit before the merge, carry over one small change that had been made on top of the older source since (settings nudges carrying their data), run the room tests, deploy, and confirm by hand that a subprotocol handshake opened.
This was also a security regression, not just an outage. The older source had no per-device budget on the inbox and no refusal logs. For four hours, the protections from the 15 September assessment were not in production.
Deploy, prove it, or roll back.
Cloudflare's December post lists what it wants for configuration changes: “health validation and quick rollback capabilities among other things.” That is the right idea at any size. The same evening, we made npm run deploy in the worker folder run scripts/deploy-gate.mjs, and made that the only supported way to ship the worker. Its header says why: “This script makes that impossible to repeat.”
Step 4 is the one that would have caught 26 September in seconds. It signs throwaway tickets and opens four sockets against the live worker: the account room and a session room, each with the desktop handshake and the phone handshake. A desktop socket only counts as passing if it opens, gets a state frame, and the worker echoed ow-ticket.
The same probe also runs on its own. A small script calls the gate with --probe-only every five minutes on our Mac. On failure it raises a macOS notification titled “Luminair rooms worker is broken” and leaves a FAILED status file that the next working session sees.
Two failed dials, then try the other door.
A gate protects the next deploy. It does not help the Macs that are already out there if something else breaks the handshake. So the socket code in the app changed too, in the same commit.
First, a socket that dies before it opens is now reported. The account room logs “dial failed” with the close code, at most once a minute, so a dead worker leaves a trace instead of four silent hours.
Second, the worker accepts the ticket two ways, and the app now knows both. If two dials in a row reach the worker but never open, the app flips to the other form, the ?t= URL the phones use, and keeps whichever form last worked until it fails twice. When that happens the log line says so plainly: “via query-ticket fallback (subprotocol handshake failing)”. It is a deliberate trade: for a while the ticket travels the less private way, and the reason is written down where we will see it.
Checked, and not claimed.
What this post does not claim
- How many users noticed. The symptom was late sync, and we did not measure who saw it.
- That the two relay lineages are reconciled. They are not yet; the gate names the four tests that prove it.
- That this was Cloudflare's fault. Workers did exactly what we asked. The comparison is about the shape of the mistake, not the scale.
Sources
- Cloudflare Blog · 18 November 2025Cloudflare outage on November 18, 2025
- Cloudflare Blog · 5 December 2025Cloudflare outage on December 5, 2025
- Hacker News · November 2025Cloudflare outage on November 18, 2025 post mortem
Your sessions, on every device
Start a session on your Mac and follow it from your phone.