The short version
- Open-weight models such as Qwen3-Coder, Kimi K2 and gpt-oss narrowed the gap with closed ones in 2025, and tools like Ollama and LM Studio made running them locally a download.
- Luminair treats a local model as one more engine. It lands in the same picker, runs through the same runner and streams the same events as Claude or Codex.
- For Ollama, Luminair can install the runtime, download curated models with a real progress bar, and lists anything you pulled yourself. It runs the server in single-model mode so a second request queues instead of loading a second 20 GB copy.
- For LM Studio, Luminair installs nothing. It reads the models LM Studio already has and talks to its local server on port 1234.
- Open weights do not always mean local. Some of the biggest open models are far too large for a laptop, and Luminair says so rather than pretending.
The gap got small.
For a long time, “open model” meant “a year behind”. Mid 2025 changed the tone. In July, Alibaba's Qwen team introduced Qwen3-Coder, “a 480B-parameter Mixture-of-Experts model with 35B active parameters”, and claimed results on agentic coding “comparable to Claude Sonnet 4”. The same month Moonshot AI released Kimi K2, with “32 billion activated parameters and 1 trillion total parameters”, and stated that “Both the code and the model weights are released under the Modified MIT License.”
In August, even OpenAI shipped open weights. Ollama's launch post for gpt-oss explained why that mattered for ordinary machines: quantizing the model's expert weights “enables the smaller model to run on systems with as little as 16GB memory, and the larger model to fit on a single 80GB GPU.”
Those numbers tell you two things at once. The best open models are genuinely strong. And the very biggest of them do not fit on a laptop. What does fit is a tier of roughly 20 to 30 billion parameter models, around 18 to 20 GB on disk, that load on a Mac with enough memory. They need no key, no bill per token and no network once the weights are down.
That is how Luminair treats them. You might send a quick refactor or a private file to a local model, and the hard architecture question to a frontier one, in the same window and even in the same thread.
Ollama and LM Studio, side by side.
A model file on disk does nothing by itself. It needs a runtime: a program that loads the weights into memory and answers requests. On a Mac, the two most common are Ollama, a background server with a command line, and LM Studio, a desktop app with a built-in local server. Both expose an HTTP API on your own machine.
Luminair has an engine file for each, plus dedicated lanes for a few curated models that run on Ollama. They all appear under the Local tab in Settings › Models, headed “On this computer”, next to the Cloud tab.
| Lane | Runs on | Model | On disk |
|---|---|---|---|
| Ollama | qwen3.8:27b | 18 GB | |
| Ollama | deepseek-r1:32b | 19.9 GB | |
| Ollama | glm-4.7-flash | 19.0 GB | |
| Ollama | muse-glimmer | ~18 GB | |
| Ollama | anything you pulled | yours | |
| LM Studio | anything in its folder | yours |
The curated sizes were chosen, not defaulted. The DeepSeek file explains why it pins the 32B model: the 70B one is 42.5 GB, “too tight next to everything else on a 64 GB Mac”, and the bare latest tag is a 5.2 GB distill “not worth a picker slot”.
Invisible, on purpose.
The goal for the Ollama lanes was that you never have to know a background server exists. Four pieces of code make that true.
1 · Install and download with real progress
If Ollama is missing when you press Download on a local model, Luminair installs it: from Ollama's app bundle on macOS, through Windows Package Manager on Windows. Then it runs ollama pull and parses its output into a progress bar with size, speed and time left. Ollama downloads a model in layers, so the parser follows the biggest layer it has seen, which is the weights. A download that stops resumes where it was when you press Download again.
Whether a model is present is read from Ollama's manifest file on disk, not from a process exit code. Deleting works the same way: Luminair asks ollama rm to free the weights (removing files by hand would leave gigabytes of orphaned layers) and then checks the manifest is gone.
2 · Anything you pulled appears
The Ollama engine is the runtime, not a vendor. It lists every model in Ollama's manifest folder, recomputed each time the picker looks, so a model you fetched in the terminal with ollama pull shows up without restarting Luminair. It skips models that cannot chat (embedding, reranking and speech models), and it skips the models a dedicated lane already owns, so each appears once, under its own vendor. Its ids start with ollama/, so they can never collide with another engine's.
3 · One model in memory, not two
This one mattered most. Ollama's defaults let it load a second copy of a model when a request arrives while the first copy is busy. On a Mac already short of memory, that second 20 GB load wedged, and the turn hung at zero tokens. We saw it twice before we understood it.
The code comment calls it “the whole fix”: one loaded model, one request in flight, and the model kept warm for 30 minutes. A second turn queues on the copy that is already in memory. If the server is not running, Luminair starts it in this mode and waits up to about 20 seconds for it to answer, showing “Starting the local model engine…” so the session never looks frozen.
4 · Local models get tools
A plain chat request to Ollama has no tools, which made local models the one lane that could not read or write the Luminair Journal. Ollama's chat API accepts the same function-calling format as the cloud APIs, so the launcher sends six Journal tools (add, append, read, search, list, remove), runs the calls the model makes, and loops for up to four rounds per turn. Whether a given model uses tools well is up to the model.
Borrow what is already there.
LM Studio is a separate app, and Luminair does not install it. If it is not on your Mac, the engine says so in one sentence: “LM Studio is not installed. Get it from lmstudio.ai, load a model, and start its local server.” No models are offered that cannot run.
When it is installed, Luminair reads LM Studio's models folder directly. LM Studio stores models as publisher/repo/file, and the id it serves over its API is publisher/repo, so that is what the picker lists, with the same filter for non-chat models. Reading the folder works whether the server is up or not.
Turns go to LM Studio's OpenAI-compatible endpoint on 127.0.0.1:1234. Luminair reuses the launcher it wrote for another OpenAI-style API: it keeps each thread as a JSON file and replays it every turn, trimmed to the newest 40 messages or 280,000 characters. If the server is off, you get a plain answer instead of a stack trace: “LM Studio is not serving on 127.0.0.1:1234. Open LM Studio, load a model, and turn its local server on.”
Honest about size.
Open weights are a licence, not a guarantee that a model runs on your desk. Two examples from the code.
Kimi K2 is open, but a trillion total parameters is not a laptop model. In Luminair, Kimi runs through Moonshot's own Kimi Code command line tool on a subscription sign-in, in the cloud. Newer GLM releases appear on Ollama only as cloud tags with no local weights, so the GLM lane pins glm-4.7-flash, which the file notes is described as “the strongest model in the 30B class”. Listing a cloud tag under “On this computer” would have been a local model that cannot run locally.
The same honesty applies to routing. Luminair's decision models have a Harness policy control called Prefer local working models. Its own description is careful: it keeps a selected local model or prefers an eligible local one, and “Jev still uses Cloudflare.” Choosing a local working model does not make every part of the app offline, and the app does not claim it does.
Three ways to go local.
- 1Open Settings › Models and choose the Local tab. Press Download on Qwen, DeepSeek or GLM. Luminair installs Ollama first if it needs to.
- 2Already use Ollama? Run
ollama pullin your terminal. The model appears under Ollama in the picker, labelled “· on this Mac”. - 3Use LM Studio? Download a model there and turn on its local server. Its models appear under LM Studio.
Then pick the model for a session, or switch to it mid-thread from the model menu to get a local second opinion on the same conversation.
Checked, and not claimed.
What this post does not claim
- How fast any local model runs on your Mac, or how its answers compare to a frontier model. We have not published benchmarks.
- That a local model uses tools well. Luminair offers the tools; the model decides.
- That the vendor benchmark figures quoted above hold for the smaller models you can run locally. They describe the flagship releases.
Sources
- Qwen Team · 22 July 2025Qwen3-Coder: Agentic Coding in the World
- Moonshot AI · GitHub · July 2025Kimi K2: Open Agentic Intelligence
- Ollama Blog · 5 August 2025OpenAI gpt-oss
For how Luminair picks between all these models, read The right model for every turn.
Run a model on your own Mac
Download one from Settings › Models, and keep your frontier engines one click away.