← Blog Architecture 26 September 2026 9 min read

The deploy that closed every Mac's socket for four hours.

Cloudflare's two big outages of late 2025 had the same shape: one change reached the whole network at once. In September 2026 we did a small version of that to ourselves, on Cloudflare Workers. Nothing crashed, nothing logged, and phones kept working. Here is what happened and what now stands in front of every deploy.

Fig 01  Three days from merge to outageTimes from our incident note · UTC on 26 Sep
MAC SOCKETS PHONE SOCKETS 23 Sep · merge resolution worker source swapped for the older lineage cc2eaa3c · nothing deployed yet 26 Sep 14:51 · wrangler deploy 18:54 · hardened source back every handshake refused, close 1006, no log line connected the whole time (ticket in the URL query) WHAT A USER SAW “The phone syncs, but minutes late.” SAME EVENING deploy gate · live probe · client fallback committed
Dates from git (commits cc2eaa3c and bbaf729f) and times from our incident note of 26 September. Lane lengths are not to scale. Mac sockets means both the account room and every session room.

The short version

  1. Luminair keeps your Macs and phones in sync through a small rooms worker on Cloudflare. Each device holds a live socket to it.
  2. On 23 September a merge-resolution commit silently replaced the worker's source with an older version that could not read the way Macs send their ticket.
  3. On 26 September a routine deploy shipped it. For about four hours, every Mac socket was refused. Phones kept working, and the app logged nothing.
  4. Now the only supported deploy runs a gate: a source check, the tests, a live handshake probe, and an automatic rollback if the probe fails. The app also falls back to the other handshake by itself.
02One change, everywhere

Two outages with the same shape.

On 18 November 2025, Cloudflare's network began failing at 11:20 UTC. A permissions change in a database made a Bot Management “feature file” double in size, and “the larger-than-expected feature file was then propagated to all the machines that make up our network.” Cloudflare called it its worst outage since 2019. Among the follow-ups, it promised to harden “ingestion of Cloudflare-generated configuration files in the same way we would for user-generated input”.

Two and a half weeks later, on 5 December, it happened again, for about 25 minutes and roughly 28% of Cloudflare's HTTP traffic. The trigger this time was a change made through the global configuration system, which, in Cloudflare's words, “does not perform gradual rollouts, but rather propagates changes within seconds to the entire fleet”. The post draws the lesson itself: “In both cases, a deployment to help mitigate a security issue for our customers propagated to our entire network.”

The November post-mortem was one of the most read engineering write-ups of the year, with more than 1,400 points on Hacker News. The reason it landed is that the pattern is universal. Every team has some path where one change goes straight to everyone.

Not a bad change. A change with no gradual path and no check on the far side.

We are nowhere near Cloudflare's scale. But in September 2026 we found one of those paths in our own setup, and it ran on their platform.

03What the worker does

A switchboard for your devices.

Luminair's desktop app runs the sessions. Your phone, your other Macs and your teammates watch and steer them. The live part of that runs through one Cloudflare Worker with two kinds of Durable Object: an account room, one per user, that every Mac and phone on the account joins, and a session room for each shared session. Prompts, replies, presence and queue changes travel as frames through those rooms.

To join a room, a device shows a ticket: a short, signed pass minted by our backend after it checks who you are. The worker verifies the signature and lets the socket in.

Where that ticket travels matters. Phones put it in the URL, as ?t=…. On 15 September, a security assessment of the computer-to-computer relay flagged exactly that as finding F-09: a bearer ticket in the URL query string, with request logging switched on. So the desktop app moved its ticket into the WebSocket subprotocol header instead. It offers two protocols, ow-ticket and ow-t.<ticket>, and a hardened worker reads the ticket from there and answers with ow-ticket.

Fig 02  The same Mac, two workersFrom the worker source
MAC Opens socket ow-ticket, ow-t.<ticket> HARDENED WORKER · 15 SEP ticketFromProtocols(header) answers 101, protocol ow-ticket also: rate budgets, xfer frames Socket open frames flow OLDER LINEAGE · SHIPPED 26 SEP reads ?t= from the URL only finds none: 401 Bad ticket header ignored Close 1006 before open, so never logged
Filled dot: the socket opens. Hollow dot: it dies in the handshake. The older worker checks the ticket from url.searchParams.get('t') and nothing else, which is why phones, which use ?t=, were unaffected.
04How the source changed

A merge that looked routine.

In mid-September two lines of work touched the worker at the same time. One was the hardened relay from the 15 September assessment: ticket in the subprotocol, per-device message budgets, computer-to-computer transfer frames, structured refusal logs. The other was an older “relay phase one” lineage, which had its own feature, computers following a session.

They were never fully reconciled. Each version fails the other's tests: the hardened source fails four “following” tests in the account-feeds file, and the older one fails the hardening tests. Production had been running the hardened one.

On 23 September, a commit titled “Resolve merged relay and Scrapbook branches” settled a conflict in desktop/rooms/src/index.js by taking the older side. The diff removed ticketFromProtocols, the subprotocol echo, the rate budget and the transfer frames. Nobody deployed the worker that day, so nothing broke. The swapped source simply sat in the repository, looking like any other file.

Three days later, someone ran wrangler deploy from that folder for an unrelated change. Wrangler did exactly what it was told. It uploaded the source it found to every Cloudflare location at once. That is the whole point of Workers, and on this day it was the problem.

05Four quiet hours

Nothing crashed. Nothing logged.

From 14:51 UTC to 18:54 UTC on 26 September, every Mac's account-room socket and every session-room socket failed its handshake. Three things made it hard to see.

01

The failure was before “open”

The shared socket code logged when a socket opened and when an open socket closed. A socket refused during the handshake never opened, so it produced neither line. It just retried, with backoff, forever.

02

Phones still worked

Phones send the ticket in the URL, which the older worker understood. Their own sockets stayed up the whole time.

03

A fallback hid the rest

With the live path down, phones still caught up through the slower Firestore polling path. So the symptom was not “sync is broken”. It was “the phone syncs, but minutes late”.

It was found by chasing that lag. The fix was to restore the hardened source from the commit before the merge, carry over one small change that had been made on top of the older source since (settings nudges carrying their data), run the room tests, deploy, and confirm by hand that a subprotocol handshake opened.

This was also a security regression, not just an outage. The older source had no per-device budget on the inbox and no refusal logs. For four hours, the protections from the 15 September assessment were not in production.

06The deploy gate

Deploy, prove it, or roll back.

Cloudflare's December post lists what it wants for configuration changes: “health validation and quick rollback capabilities among other things.” That is the right idea at any size. The same evening, we made npm run deploy in the worker folder run scripts/deploy-gate.mjs, and made that the only supported way to ship the worker. Its header says why: “This script makes that impossible to repeat.”

Fig 03  What stands in front of wrangler deployFrom deploy-gate.mjs
01 Static guard 5 features must be in the source 02 Tests every file green, 4 known failures 03 Deploy after noting the live version 04 Live probe 2 rooms × 2 handshakes Shipped any probe fails 05 · Roll back stop Steps 01 and 02 stop before anything is deployed.
The five features the static guard looks for are the subprotocol ticket parser, the header that hands the chosen protocol to the room, the protocol echo on the 101 response, the per-device budget, and transfer frames. If the probe cannot read the room secret, the gate refuses to run rather than skip the check.

Step 4 is the one that would have caught 26 September in seconds. It signs throwaway tickets and opens four sockets against the live worker: the account room and a session room, each with the desktop handshake and the phone handshake. A desktop socket only counts as passing if it opens, gets a state frame, and the worker echoed ow-ticket.

The same probe also runs on its own. A small script calls the gate with --probe-only every five minutes on our Mac. On failure it raises a macOS notification titled “Luminair rooms worker is broken” and leaves a FAILED status file that the next working session sees.

07The client heals itself

Two failed dials, then try the other door.

A gate protects the next deploy. It does not help the Macs that are already out there if something else breaks the handshake. So the socket code in the app changed too, in the same commit.

First, a socket that dies before it opens is now reported. The account room logs “dial failed” with the close code, at most once a minute, so a dead worker leaves a trace instead of four silent hours.

Second, the worker accepts the ticket two ways, and the app now knows both. If two dials in a row reach the worker but never open, the app flips to the other form, the ?t= URL the phones use, and keeps whichever form last worked until it fails twice. When that happens the log line says so plainly: “via query-ticket fallback (subprotocol handshake failing)”. It is a deliberate trade: for a while the ticket travels the less private way, and the reason is written down where we will see it.

The lesson we tookA deploy is not finished when the upload succeeds. It is finished when a real client, using the real handshake, has been let in. Anything short of that is a hope.
08What we checked

Checked, and not claimed.

40/44
Room worker tests pass today. The 4 failures are the known account-feeds “following” tests the gate allows
4
Live handshakes per probe: two rooms, desktop and phone form
5 min
Interval of the standing probe on our Mac
~4 h
14:51 to 18:54 UTC on 26 September, Mac sockets refused

What this post does not claim

  • How many users noticed. The symptom was late sync, and we did not measure who saw it.
  • That the two relay lineages are reconciled. They are not yet; the gate names the four tests that prove it.
  • That this was Cloudflare's fault. Workers did exactly what we asked. The comparison is about the shape of the mistake, not the scale.

Sources

Your sessions, on every device

Start a session on your Mac and follow it from your phone.

Download Luminair →