← Blog Building Luminair 26 September 2026 9 min read

What did the agents actually do today?

The big studies disagree on whether AI makes developers faster. We cannot settle that. What we can do is keep an honest record of one day: what the agents checked, what did not finish, and where the tokens went.

Fig 01  One day, three recordsSchematic · local time
WHAT IT CHECKED Routine check, every 15 min WHAT HAPPENED Tracked, finished, failed WHERE TOKENS WENT One record per finished turn flagged: possibly stalled tracked finished verified job failed same job 09:0011:0013:0015:0017:00 YOUR DAILY RECAP 1 task did not finish Tried 2 times counted from the log, not from model prose Hollow dots are routine checks that found nothing. The ledger lane never calls a model; the recap only asks one for a short summary and ideas.
Schematic of one day. The rules are real: a routine check every 15 minutes while Solace is on, one ledger record per finished turn, and a recap after 17:00 whose counts come from the log. The events and their times are illustrative.

The short version

  1. Developers overwhelmingly say AI makes them more productive. The best controlled study found the opposite in 2025, then said in 2026 that it could no longer measure the effect cleanly.
  2. We do not try to answer that. We answer a smaller question every day: what did the agents actually do, including the checks that found nothing.
  3. Luminair's Solace logs a routine check every 15 minutes while it is on, and writes a daily recap whose counts come from that log, never from the model's prose.
  4. A separate token ledger records every finished turn with no model call, and counts two habits worth changing: light tasks on premium models, and Max or Extra-high thinking.
02The paradox

Faster, says the survey. Slower, said the stopwatch.

In September 2025 Google's DORA team published its State of AI-assisted Software Development report. The announcement put the headline plainly: “90% of survey respondents report using AI at work. More than 80% believe it has increased their productivity.” In the same paragraph, 30% reported little or no trust in the code the AI wrote.

The report's central line is the useful one: “AI doesn't fix a team; it amplifies what's already there.” It also found that AI adoption “does continue to have a negative relationship with software delivery stability”, and that without “strong automated testing, mature version control practices, and fast feedback loops, an increase in change volume leads to instability.”

Then there is METR, which ran a randomized study with experienced open-source developers. Its early 2025 result, restated in a February 2026 update, was that “the use of AI causes tasks to take 19% longer”. The update is more interesting than the number. Developers now refuse to work without AI, even at $50 an hour, and “30% to 50% of developers told us that they were choosing not to submit some tasks because they did not want to do them without AI.” METR concluded its new data “is only very weak evidence” for the size of any change, and is redesigning the study.

If the researchers cannot measure it cleanly, your gut feeling about last Tuesday is not a measurement either.

We cannot tell you whether running many agents makes you faster. What we can give you is DORA's “fast feedback loop” at the scale of one person and one day: a record of what the agents did, what broke, and what it cost, written so you can act on it tomorrow morning.

03What it checked

A quiet day should not look empty.

Solace is the part of Luminair that tracks every “I'll report back” promise across your sessions until it is done, verified and confirmed. It is event-driven: it wakes when a session finishes, a background job ends, or a verifier returns.

That design had an awkward side effect, recorded in the commit of 27 July 2026. Between events, its watching left no trace, so “an uneventful day looked empty/useless”. A log that only records changes cannot tell a day where nothing broke from a day where nothing was watched.

The fix was a patrol. While Solace is on, it runs a routine check shortly after launch and then every 15 minutes, and logs what it reviewed and what it found, including the honest “nothing needed doing”. If a tracked item has been in progress for more than 20 minutes with no update and no background job behind it, the check flags it as possibly stalled.

Solace · activity dropdown
MC ReportsClear
Today
Checked ×12: Routine checkWatched for new “report-back” work and re-checked the ledger. Nothing tracked, nothing to do.2:15 to 5:00 PM
Verify failed: the verifier's one-line reason2:02 PM
Checked: Routine checkReviewed 3 tracked items (1 open, 1 working, 1 awaiting verify). Flagged 1 as possibly stalled.11:45 AM
Advanced: Advanced11:20 AM
Tracked: Tracked9:40 AM
Startup check: On launchReviewed 3 tracked items after startup, all consistent, nothing to recover.9:02 AM
a changea routine checkneeds a human
Drawn from the labels in the app's code (CTL_LOG_LABEL and controllerPatrol in desktop/renderer.js). Task names are blanked; times and counts are illustrative. Identical checks on the same day collapse into one row with a ×N count and a time range, and the row expands to show every check's time.

The collapse also shapes the recap: the model sees “Routine check ×12”, not twelve duplicate lines, which is cleaner and costs fewer tokens.

04The daily recap

Counted by code, told by a model.

Once a day, after 17:00 local time, Solace writes a recap from that log. The split of work is deliberate. Code counts; the model only writes a short summary and suggests improvements.

  • The counts are derived, not remembered. Every time the card opens, its numbers are recomputed from the day's log. Only the model's summary stays frozen as written.
  • The verdict comes first and comes from the log. The source comment for the 24 September redesign says the card should answer one question, “does anything need me?”, with a verdict “computed from the log (never from the model's prose)”.
  • A quiet day never reaches the model. With nothing logged, an early version handed the model an empty log and got a greeting back. Now an empty day gets a short local note explaining what the recap measures, and no model call.
  • Suggestions are opt-in rules. The model may propose two to four improvements. Each has Turn on and No thanks. Turning one on asks for confirmation, then makes it a standing rule for future turns; anything that reads like granting tools, permissions or credentials, or weakening security, is refused before it can be adopted. Dismissed ideas are passed back to the model with an instruction not to suggest them again.

One bug shows why counting in code matters. On 24 September a failed background job was being filed under “maintenance”, so the recap said zero errors on a day with five failed builds. The fix classifies a log line by what it says, not only by its type: a line saying a job failed, or finished with a non-zero code, counts as an error.

05Plain reasons

Nobody should need to know what 143 means.

Two days later, on 26 September, the card was rewritten again for people who do not read exit codes. The note in the source is blunt: the old card was “horrible for general users”. Three changes came out of it.

First, one row per task, not per attempt. Three failed tries at the same deploy are one thing to look at, shown once with “Tried 3 times”. Second, the model's summary is told to write for a non-technical reader: say “the assistant” rather than internal names, never mention exit codes or log fields, and end with the one thing worth doing next. Third, exit codes became reasons:

Fig 02  From exit code to reasonReal mapping · desktop/renderer.js
WasNow
exit 143, 137, 130Stopped when the app closed or restarted
exit 65The app build did not go through
exit 127, 126A tool it needed was missing
exit 124Took too long and was stopped
verify: failFinished, but the result did not check out
interruptedStopped before it finished
any other codeStopped with an error
The complete list from the needWhy function. The raw code is not thrown away: hovering a row shows “Technical detail: exit code 143” for whoever needs it.

The rest of the card follows the same idea. The small counters read Finished, Double-checked, Fixed itself and Cleanups, and every non-zero number opens the exact log lines behind it.

Your daily recap
Your daily recap
!
1 task did not finish2 tries in total · Tap a task to see what happened
sending the app to the iPhoneStopped when the app closed or restarted
Tried 2 times · last
Finished
Double-checked
0Fixed itself
0Cleanups
What happened today

A short summary written by the model, collapsed until you ask for it.

Ideas to make tomorrow smoother
Turn onNo thanks
Drawn from the card's labels in the source. “Sending the app to the iPhone” is the example task name the summary prompt itself uses. Numbers are blanked because they depend on your day.
06Where the tokens went

A ledger that costs nothing to keep.

The recap tells you what happened. The second record tells you what it cost. Since 4 August 2026, every finished turn drops one compact record into a local token ledger. The design rule is written into the code in capitals: no model call, because this must stay free.

Each record lands in a bucket for your local day, split three ways: by model, by thinking level and by session. The ledger keeps 14 days per account and prunes on every write. It is fed from both places a turn can finish: the desktop app and turns relayed from your phone.

On top of the totals, the ledger keeps two coaching counters. Each is a plain rule, not a judgement by another model:

Fig 03  How a finished turn is countedReal rules · desktop/main.js
EVERY TURN finishes ALWAYS Add tokens to today · by model · by thinking level · by session COUNTER 1 · LIGHT TASK, PREMIUM MODEL ✓ prompt shorter than 240 characters ✓ thinking set below High ✓ model pricier than that engine's cheapest all three COUNTER 2 · DEEP THINKING thinking level is Max or Extra-high
The rules in _turnIsLight, isPremiumModel and recordTokenSpend. “Premium” is read from each engine's own model list, so a new engine's expensive models count as premium the day it lands, with no change to the report.

The two counters then drive one sentence and a couple of tips. If less than 12% of the day's tokens went to those two habits, the card calls it a lean day; under 35%, balanced; above that, a rich day with room to trim. The tips are ranked by how many tokens each would move, and read like Swap quick questions onto Sonnet 5 and Ease off Max thinking. When neither counter fires, the card says so and suggests nothing.

The first version counted tokens only, on purpose: the commit says “never invented dollar prices”. Dollars arrived on 6 September with the Burn meter, and the rule survived in a stricter form. A Claude turn that reports its own cost is recorded as reported. Anything else is priced from a published list-price book and marked est. A model missing from the book stays unpriced rather than getting a guess, and turns that reported no usage at all are counted and shown, so a partial total never passes for the whole day. The panel's footnote is honest about the rest: “Subscription figures are API equivalents, not your subscription bill.”

Why deterministicA report whose numbers come from a model can be wrong in ways you cannot check. Both records are plain arithmetic over things that already happened. The only model in the loop writes the recap's summary, and the card is built so you never need to read it.
07Find it in the app

Three places, one panel.

  1. 1Click the small arrow next to the Solace chip. The MC Reports dropdown lists every step it took today, routine checks included. Click a merged row to see each check's time.
  2. 2Open Solace and choose the Spend tab. The Burn meter shows today's dollars by model, project and session, and Your token diet · today shows the per-model bars, the day's verdict and the swaps.
  3. 3The daily recap card appears after 17:00. For now it is switched on only for Luminair's own admin accounts: we are running it on our own work first.

Solace is part of the Pro+ plan. The pricing page has the details.

08What we checked

Checked, and not claimed.

21/21
Solace audit and behaviour tests pass, including the ledger's day buckets and the Burn meter's weekly totals
15
Minutes between routine checks while Solace is on
14
Days of token history kept per account, pruned on every write
0
Model calls made to keep the token ledger

What this post does not claim

  • That Luminair makes you more productive. We have not measured that, and the studies above show how hard it is to measure.
  • That the coaching counters are always right. A short prompt can still be hard work; the counters flag a pattern, not a mistake.
  • That the recap is available to every user today. It is limited to admin accounts while we run it on ourselves.
  • That dollar figures match your bill. Subscription turns are shown as API equivalents, and estimates are marked.

Sources

See your own day

Open Solace, choose Spend, and see where today's tokens went.

Download Luminair →