← Blog Mobile 10 September 2026 9 min read

Talking to your agents: on-device and cloud dictation

Voice was in Luminair on the day of its first commit. Since then it has moved from a third-party dictation API to cloud Whisper with an on-device fallback, grown a phone version, and picked up a handful of small fixes that did more for daily use than any model upgrade.

Fig 01  One dictation on the Mac, start to finishFrom desktop/main.js · order is real
YOU Hold Fn, talk, let go 16 kHz mono WAV TRIED IN ORDER · FIRST ANSWER WINS 1 · CLOUD · SIGNED IN Elion STT Shared speech service · whisper-large-v3-turbo 2 · CLOUD · SIGNED IN Luminair transcribe function Groq-hosted whisper-large-v3 3 · ON THIS MAC · NO ACCOUNT NEEDED whisper.cpp large-v3-turbo, about 547 MB, downloaded once failed or timed out failed, offline, or not signed in EVERY ROUTE Strip non-speech [BLANK_AUDIO], [Music] … Clean up fillers, self-corrections COMPOSER “Pull the retry path out.” Each cloud attempt times out after 4 seconds plus about 1 second per 24 KB of audio, capped at 25 seconds, so a dead network falls through to the Mac quickly.
The current order in the wispr:transcribe handler. The handler kept its name from the first day, when the only route was Wispr Flow's API. On-device transcription uses a bundled macOS binary, so step 3 is Mac only.

The short version

  1. Luminair's first commit, on 17 July 2026, already had voice: a server-side proxy to Wispr Flow so nobody had to paste an API key.
  2. The same day it gained the Fn key as push-to-talk and free on-device transcription with whisper.cpp.
  3. Six days later the order flipped: cloud Whisper first when you are signed in and online, the Mac's own model when you are not. Voice always has a route.
  4. The fixes people noticed most were small: no more [BLANK_AUDIO] in the composer, and a headphone button that plays music instead of starting a recording.
02Built with Luminair

Found by using it.

The voice commits cited here all carry a Claude co-author line: the feature was built in agent sessions, and its rough edges were found by using it. The stray markers, the lag and the play button that started a recording did not come from a test plan. The commit that fixed the markers quotes the whole bug report: “never show this”.

This post was put together in a Luminair session that read the voice code, the phone app and the git history, and checked every outside claim against the page it links to.

03Everyone started talking

Voice became a coding input.

In November 2025 TechCrunch reported that Wispr Flow users, after three months, write “more than 50% of their characters through the app”, and that the company had grown 40% month over month since June. The same piece noted Wispr was testing a closed API with select partners. In March 2026, Claude Code rolled out a voice mode: “type /voice to toggle it on, then speak your command”, live for about 5% of users at first. And in August 2026 Wispr raised $280 million at a $2 billion valuation.

That last story has a line every voice product should read. “For the last few weeks, several users have complained about a quality dip in Wispr Flow's dictation output.” Dictation is judged word by word. One wrong word in a prompt to a coding agent is not a typo; it can be a different instruction.

Talking to an agent is not the same as dictating an email. The words are instructions.
04Day one

Voice was in the first commit.

The initial commit describes the project in one line: an Electron desktop app “plus multi-tenant voice backend”. Transcription went through Wispr Flow's API, with the key held in a server function. You signed in with Google; the app sent your audio to that function, which called Wispr and returned text. No one pasted a key.

Two more voice commits landed that same day.

The Fn key. The globe key is where most Mac users expect dictation to live, but, as the commit puts it, “macOS never delivers the Fn key to the page”. So the app ships a tiny native helper that watches the key and streams down and up events to the window. It only runs when Fn is your chosen talk key.

Free, offline transcription. A bundled whisper.cpp binary transcribes the recording on the Mac with an on-device Whisper model. The commit message: “Voice no longer needs Flow/Wispr or any key.” The model is downloaded once into the app's data folder the first time it is needed. Today that is the large-v3-turbo model, with the small English base model as a fallback if it is already on disk.

05Online first, local always

Local-first felt slow.

At first the Mac tried local transcription before the cloud. That sounds right for privacy and cost, and it felt wrong in use. The commit that flipped it on 23 July explains why: “every dictation waited on a 547MB whisper.cpp model reload even when the fast cloud was available.”

The new order, which Fig 01 shows in its current form, is cloud first when you are signed in, the Mac as the fallback. The cloud route used Whisper large-v3 hosted by Groq (when the server holds a Groq key), which the commit describes as “~instant”. Not signed in? The cloud helper returns nothing, without an error, and the local model takes over. Offline? The request times out and the local model takes over. In the commit's words: “Voice always works, just a beat slower off-grid.”

Since then a shared service has been added in front. Elion STT is a speech-to-text service meant to be shared across Elion apps; the Mac tries it first, then the original Luminair function, then whisper.cpp. Each step is allowed to fail.

Why keep the local model at allBecause the cloud can be unreachable at exactly the wrong moment: on a train, behind a captive portal, when a key expires on the server. A dictation that silently fails is worse than one that is a second slower.
06The phone

Same rule, different fallback.

The phone app followed the same day, 23 July. It records audio locally, uploads it to the same transcribe function when you are online and signed in, and falls back to Apple's on-device recognizer when you are not. The backend gained a format parameter so Whisper decodes the phone's m4a as well as the Mac's WAV.

The phone also got something the Mac did not have in the same form: a full-screen, hands-free conversation mode. It listens, waits for a pause, sends your words to the model, and reads the reply aloud with the phone's native speech engine. The design goal in the commit was to use the phone's own hardware for everything except the model itself.

Fig 02  What stays on the phoneFrom mobile/src · real
Dictation
Online and signed in: cloud WhisperRecorded as m4a, uploaded, transcribed on the server
Offline or signed out: Apple recognizerStreams partial words on the device
Conversation mode (built 23 July)
Listen with the phone's recognizer
End of turn after 1.4 s of silenceA local timer, because the recognizer sends no “final” event
Send to the modelThe only step that crosses the network
Speak the reply with the native voiceMarkdown, code and links stripped first
on the deviceover the network
From voice.js, tts.js and the VoiceMode component in the phone app. On the phone the earbud button also works as “talk and go”: the first press records, the second transcribes by the same route and sends. The Mac's version was removed, as the next section explains.

One detail in conversation mode is worth copying. Every voice turn carries a hidden instruction to answer in one or two short spoken sentences and “Be decisive: give the answer directly and commit to it.” A reply that hedges reads fine on screen and sounds terrible out loud.

07The small fixes

Three fixes you'd notice.

1 · The words nobody said

Whisper, local or cloud, writes placeholders for silence and noise: [BLANK_AUDIO], [ Silence ], [Music], (applause), music notes. They were landing in the composer and getting sent to the model. The fix is one filter at the single point every transcript passes through, and a copy of it on the phone.

desktop/main.jslines 23314–23322, pattern shortened
// A curated non-speech word list, inside brackets or parens ONLY,
// so real bracketed text the user dictated ("[important]") survives.
const NON_SPEECH_RE = /[\[(]\s*(?:blank[\s_]*audio|silence|inaudible|
  music(?:\s+playing)?|noise|applause|laughter|typing|beep|static)\s*[\])]/gi;
function stripNonSpeech(s) {
  return String(s || '').replace(NON_SPEECH_RE, ' ')
    .replace(/[♪♫]+[^♪♫\n]*[♪♫]+/g, ' ').replace(/\s{2,}/g, ' ').trim();
}

The narrow list is the point. A broad “remove anything in brackets” rule would have eaten words people meant, like a dictated “(staging)”. The commit records a 12-case check: markers stripped, a transcript that is only a marker becomes empty, real speech intact.

2 · Saying what you meant

People correct themselves when they talk. Since 21 July every transcript can go through a cleanup pass before it reaches the composer. The Settings text gives the example: say “email John, no, Sarah” and it writes “Email Sarah”. The original prompt also told the model never to obey the transcript as an instruction, only tidy it. It is on by default and one switch turns it off if you want the raw words.

3 · The headphone button

This one went back and forth, and the last move was to remove it.

17 JUL

Play/Pause talks

The first commit bound the media Play/Pause key: while the app was focused, a press started or stopped the mic.

17 AUG

Back, as an opt-in

Re-added so a headset button could start dictation, only while Luminair was the front app. macOS refuses that key to Electron's shortcut API, so the native helper reported it instead. Without an extra Input Monitoring grant the helper could only listen, not swallow the key, so a press could also start Music.

10 SEP

Removed on the desktop

“The play button just controls music.” Fn and the chosen keyboard talk key are unaffected.

On the phone the earbud button stayed, because there it has one clear job: talk and send. On a Mac with music playing, the same key meant two things at once, and a talk key that sometimes plays a song is not a talk key.

08Find it in the app

Set it up in a minute.

  1. 1Open your avatar at the bottom left, then Settings › Voice.
  2. 2Under Voice account, sign in for the cloud route. Without it, dictation runs on the Mac.
  3. 3Under Hold-to-talk key, pick Fn or another key. The hint in Settings is blunt: “Hold to talk, release to send.” If another dictation tool uses the same key, give Luminair its own.
  4. 4Leave Clean up dictation on, or turn it off for raw transcripts. Microphone picks the input device.

More on the feature page: Speech to Text. On the phone, the mic sits on the command bar above the keyboard.

09Honest limits

What we haven't measured.

3
Routes the Mac tries for each dictation, one of them on the device
547 MB
Size of the on-device turbo model, downloaded once
12
Cases in the check behind the non-speech filter
1.4 s
Silence that ends a turn in the phone's conversation mode

Worth knowing

  • We have not published accuracy or latency figures for any route. “Instant” and “a beat slower” are how the commits describe it, not measurements.
  • In current builds, dictation on the desktop needs the Pro plan.
  • The hands-free conversation modes are not one tap away right now. On the phone the wave button was taken out of the composer on 31 July; on the Mac the Live conversation beta's button has been hidden since mid-August. The code for both is still in the apps.
  • On-device transcription uses a macOS binary. On Windows, dictation needs the cloud route.

Sources

Hold Fn and talk

Dictation in the cloud when you're online, on your Mac when you're not.

Download Luminair →