The short version
- In 2025 Chinese labs released the strongest open-weight models. Moonshot's Kimi K2 Thinking, in November, was one of the headline releases.
- Luminair runs Kimi through Moonshot's own Kimi Code CLI, signed in with a Kimi Code subscription rather than an API key. DeepSeek R1, GLM 4.7 Flash and Qwen3 are downloaded and run on your computer through Ollama.
- Kimi's usage meter was written by reading the bundled CLI's own source for the endpoint its usage panel calls, and it says so in the code.
- These CLIs move fast. Within one week in September 2026, Dependabot's proposed Kimi Code upgrade went from 0.41.0 to 2.1.1. Luminair stays pinned to the 0.36.0 it verified, and stops the CLI from updating itself.
Verified live, and marked where not.
Every one of these engines was added in Luminair sessions, and their files read like lab notebooks. The Kimi engine records what was checked against the real CLI and when: the device-code sign-in on 15 August 2026, the token file's location on 7 September. It also flags, with a warning sign, the two things that still needed a real account to confirm.
That habit matters more for these models than for most. Their tools change quickly, their documentation is sometimes thinner in English, and the easiest mistake is to write code against what you assume a CLI does. This post was drafted the same way, from the engine files and the git history.
The top of the chart changed.
Simon Willison's review, 2025: The year in LLMs, gives the shift its own heading: “The year of top-ranked Chinese open weight models”. Of 2024 he writes that the Chinese labs' models “were neat models but didn't feel world-beating. This changed dramatically in 2025.” In the Artificial Analysis ranking of open-weight models he cites, as of 30 December 2025, the top five were all Chinese, and “The highest non-Chinese model in that chart is OpenAI's gpt-oss-120B (high), which comes in sixth place.”
Kimi K2 Thinking was one of the late-year highlights. On 6 November 2025 Willison described it as “a trillion parameters (MoE, 32B active)”, released under Moonshot's modified MIT licence, and asked: “Could this be the first open weight model that's competitive with the latest from OpenAI and Anthropic, especially for long-running agentic tool call sequences?” CNBC reported, citing a source, that it “cost $4.6 million to train”, and noted it was unable to independently verify that figure.
For a developer, the practical upshot was simple. Some of the best models for agentic coding were now either cheap to call or free to download. An app that wants to be “one window for every model” has to include them.
Rent the model, or own the weights.
Open weights do not mean you should always run them yourself. Kimi K2 Thinking is about 594 GB on Hugging Face, by Willison's count, far beyond a laptop. So Luminair takes two routes, shown in Fig 01.
Kimi runs in Moonshot's cloud. The engine file explains the choice: it signs in with a Kimi Code subscription instead of a pay-per-token API key, because for steady use the flat plan costs less, and it never stores a key. Willison's year review makes the same point about coding plans in general: tools like Claude Code and Codex CLI “can burn through enormous amounts of tokens”, so a monthly plan “offers a substantial discount”.
DeepSeek, GLM and Qwen run on your computer. Each engine picks one tag, and the files say why. DeepSeek's 70B tag is 42.5 GB, “too tight next to everything else on a 64 GB Mac”, and its bare tag is a 5.2 GB distill “not worth a picker slot”. GLM's newer 5.x tags on Ollama are cloud-only, “so listing them here would be a local model that cannot run locally.”
| Engine | Model in the picker | Runs | Effort levels |
|---|---|---|---|
| Kimithe plan's default model | Moonshot cloud | low · high · max | |
| DeepSeek R1 · on this Macdeepseek-r1:32b | Your computer | low to max, 5 levels | |
| GLM 4.7 Flash · on this Macglm-4.7-flash | Your computer | low to max, 5 levels | |
| Qwen3 · on this Macqwen3.8:27b | Your computer | low to max, 5 levels |
The local three share one piece of machinery, so adding the second and third was mostly a matter of choosing the tag. If Ollama is missing, the app installs it. Once the weights are on disk, an invisible local account is created, because every turn in Luminair is signed by an account, even one with no login. For the full story of the local side, see Open weights on your Mac.
A login that can succeed and fail.
Kimi's lane works like Luminair's Codex lane. The bundled CLI runs in print mode and streams one JSON object per line, and -S resumes its own thread, so the session id Luminair stores is Kimi's thread id. The CLI is a Node script rather than a native binary, so the app runs it through its own Electron runtime instead of trusting a node on your PATH.
Signing in is a device-code flow: kimi login opens the browser to Kimi's authorize page and waits. Luminair gives each Kimi account its own private home folder, so several accounts can sit side by side, and it shows you the code the CLI printed so you can check the browser shows the same one.
Two failures shaped the code. On 15 August the browser step succeeded and the CLI still exited, because Moonshot checks the Kimi Code membership afterwards; its message was “We're unable to verify your membership benefits at this time.” Luminair now keeps the CLI's last real line and shows you that, not “the window closed”. On 7 September Kimi's servers returned errors on the account and model endpoints and on chat. The token had already been saved, but the step that writes a default model had not run, and a home with no default model cannot run a turn. Login now waits for both, and retries the provisioning step up to twice, five seconds apart, without opening the browser again.
The same day found a quieter bug. Luminair had been looking for the wrong token file. The CLI writes its login to credentials/kimi-code.json; files ending in -tokens.json belong to its MCP sign-ins. So every successful Kimi login had been reported as failed.
Read the CLI, not the docs.
Every Luminair account shows how much of its plan is left. For Kimi, rather than guess at an API, we looked at the CLI: it has its own usage panel, and it ships as readable JavaScript. On 15 September the engine's meter was written by reading the bundled 0.36 source for the call that panel makes.
The code is candid about its limits. A warning in the file says it was “Written from the CLI's source; no Kimi account was signed in on this Mac to see a live payload”, and tells the next person to check it on the first real sign-in. The test suite pins the shape: a stubbed /usages and /me must become the rows Week, 5h and Booster, and fill in the plan name. For a vendor with no usage API at all, see A usage meter for a vendor that has none.
0.36 to 2.1.1 in a week.
Bundling a vendor's CLI buys you their login, their thread format and their tools. It also ties you to their release pace, and Kimi's moved quickly. Dependabot watches Luminair's dependencies and opens a pull request when a newer version appears. Here is what it proposed for @moonshot-ai/kimi-code:
A jump from 0.x to 2.x is exactly the kind of change that can move a token file, rename a flag or change a stream format, which are the three things the Kimi lane depends on. So the app stays on the version its notes were verified against, and it sets KIMI_CODE_NO_AUTO_UPDATE=1 on every call so the CLI never replaces itself under a running app. Moving to a new version means re-checking the login, the token path and the meter, not just bumping a number.
Three clicks to a new model.
- 1For the local models, open your avatar at the bottom left, then Settings › Models. Open Qwen, GLM or DeepSeek and press its download button, for example Download DeepSeek R1. Each is about 18 to 20 GB.
- 2For Kimi, add a Kimi account from the account menu and finish the device sign-in in your browser. You need a Kimi Code plan.
- 3Pick the model from the session's model menu. You can switch to it mid-thread, and back.
One limit to know: Kimi sessions run on your computer, not from the phone. The app's relay tells you so and suggests sending the turn from the desktop.
Checked, and not claimed.
What this post does not claim
- How these models compare with each other or with US models. We have not benchmarked them, and the scores quoted above are the sources' own.
- That the Kimi meter has been seen against a live account. Its code says it has not, as of when it was written.
- Which model a Kimi Code plan runs. Luminair sends no model flag and uses the plan's default.
Sources
- Simon Willison · 6 November 2025Kimi K2 Thinking
- CNBC · 6 November 2025Alibaba-backed Moonshot releases new AI model Kimi K2 Thinking
- Simon Willison · 31 December 20252025: The year in LLMs
Try a model from another lab
Download Qwen, GLM or DeepSeek, or sign in with Kimi Code, and switch mid-thread.