← Blog Models & Routing 27 September 2026 9 min read

Kimi, DeepSeek, GLM, Qwen: adding the Chinese coding models

By the end of 2025 the top of the open-weight charts was mostly Chinese. Luminair runs four of those model families next to Claude, ChatGPT and Gemini. One arrives through its vendor's own command-line tool; three are downloaded and run on your computer. Each route taught us something different.

Fig 01  Two ways inFrom desktop/engines
ROUTE A · A SUBSCRIPTION CLI MOONSHOT'S SERVERS Kimi the plan's default model internet · your Kimi Code plan BUNDLED IN THE APP kimi -p … --output-format stream-json -S <thread> ROUTE B · WEIGHTS ON YOUR COMPUTER DeepSeek R1 deepseek-r1:32b 19.9 GB GLM 4.7 Flash glm-4.7-flash 19.0 GB Qwen3 qwen3.8:27b 18 GB one download OLLAMA, INSTALLED FOR YOU IF MISSING No key, no login, no internet once the weights are down one shared launcher, the same Journal tools as every other engine One Luminair session switch between these and Claude, ChatGPT or Gemini mid-thread; the conversation is carried over
Tags and sizes as written in the engine files for DeepSeek, GLM and Qwen, each verified against the Ollama registry when it was added in August 2026. Kimi runs whichever model the Kimi Code plan sets as its default; Luminair sends no model flag.

The short version

  1. In 2025 Chinese labs released the strongest open-weight models. Moonshot's Kimi K2 Thinking, in November, was one of the headline releases.
  2. Luminair runs Kimi through Moonshot's own Kimi Code CLI, signed in with a Kimi Code subscription rather than an API key. DeepSeek R1, GLM 4.7 Flash and Qwen3 are downloaded and run on your computer through Ollama.
  3. Kimi's usage meter was written by reading the bundled CLI's own source for the endpoint its usage panel calls, and it says so in the code.
  4. These CLIs move fast. Within one week in September 2026, Dependabot's proposed Kimi Code upgrade went from 0.41.0 to 2.1.1. Luminair stays pinned to the 0.36.0 it verified, and stops the CLI from updating itself.
02Built with Luminair

Verified live, and marked where not.

Every one of these engines was added in Luminair sessions, and their files read like lab notebooks. The Kimi engine records what was checked against the real CLI and when: the device-code sign-in on 15 August 2026, the token file's location on 7 September. It also flags, with a warning sign, the two things that still needed a real account to confirm.

That habit matters more for these models than for most. Their tools change quickly, their documentation is sometimes thinner in English, and the easiest mistake is to write code against what you assume a CLI does. This post was drafted the same way, from the engine files and the git history.

03The year of Chinese open weights

The top of the chart changed.

Simon Willison's review, 2025: The year in LLMs, gives the shift its own heading: “The year of top-ranked Chinese open weight models”. Of 2024 he writes that the Chinese labs' models “were neat models but didn't feel world-beating. This changed dramatically in 2025.” In the Artificial Analysis ranking of open-weight models he cites, as of 30 December 2025, the top five were all Chinese, and “The highest non-Chinese model in that chart is OpenAI's gpt-oss-120B (high), which comes in sixth place.”

Kimi K2 Thinking was one of the late-year highlights. On 6 November 2025 Willison described it as “a trillion parameters (MoE, 32B active)”, released under Moonshot's modified MIT licence, and asked: “Could this be the first open weight model that's competitive with the latest from OpenAI and Anthropic, especially for long-running agentic tool call sequences?” CNBC reported, citing a source, that it “cost $4.6 million to train”, and noted it was unable to independently verify that figure.

For a developer, the practical upshot was simple. Some of the best models for agentic coding were now either cheap to call or free to download. An app that wants to be “one window for every model” has to include them.

Some of the best coding models became cheap to call, or free to download.
04Two ways in

Rent the model, or own the weights.

Open weights do not mean you should always run them yourself. Kimi K2 Thinking is about 594 GB on Hugging Face, by Willison's count, far beyond a laptop. So Luminair takes two routes, shown in Fig 01.

Kimi runs in Moonshot's cloud. The engine file explains the choice: it signs in with a Kimi Code subscription instead of a pay-per-token API key, because for steady use the flat plan costs less, and it never stores a key. Willison's year review makes the same point about coding plans in general: tools like Claude Code and Codex CLI “can burn through enormous amounts of tokens”, so a monthly plan “offers a substantial discount”.

DeepSeek, GLM and Qwen run on your computer. Each engine picks one tag, and the files say why. DeepSeek's 70B tag is 42.5 GB, “too tight next to everything else on a 64 GB Mac”, and its bare tag is a 5.2 GB distill “not worth a picker slot”. GLM's newer 5.x tags on Ollama are cloud-only, “so listing them here would be a local model that cannot run locally.”

Fig 02  The four enginesAs registered in the app
EngineModel in the pickerRunsEffort levels
KimiKimithe plan's default modelMoonshot cloudlow · high · max
DeepSeekDeepSeek R1 · on this Macdeepseek-r1:32bYour computerlow to max, 5 levels
GLMGLM 4.7 Flash · on this Macglm-4.7-flashYour computerlow to max, 5 levels
QwenQwen3 · on this Macqwen3.8:27bYour computerlow to max, 5 levels
Labels and effort lists from each engine's registration. Kimi's K-series takes low, high and max, so Luminair never offers medium or extra-high for it rather than rounding silently. A solid badge is a cloud model; a dashed one runs locally.

The local three share one piece of machinery, so adding the second and third was mostly a matter of choosing the tag. If Ollama is missing, the app installs it. Once the weights are on disk, an invisible local account is created, because every turn in Luminair is signed by an account, even one with no login. For the full story of the local side, see Open weights on your Mac.

05The Kimi lane

A login that can succeed and fail.

Kimi's lane works like Luminair's Codex lane. The bundled CLI runs in print mode and streams one JSON object per line, and -S resumes its own thread, so the session id Luminair stores is Kimi's thread id. The CLI is a Node script rather than a native binary, so the app runs it through its own Electron runtime instead of trusting a node on your PATH.

Signing in is a device-code flow: kimi login opens the browser to Kimi's authorize page and waits. Luminair gives each Kimi account its own private home folder, so several accounts can sit side by side, and it shows you the code the CLI printed so you can check the browser shows the same one.

Two failures shaped the code. On 15 August the browser step succeeded and the CLI still exited, because Moonshot checks the Kimi Code membership afterwards; its message was “We're unable to verify your membership benefits at this time.” Luminair now keeps the CLI's last real line and shows you that, not “the window closed”. On 7 September Kimi's servers returned errors on the account and model endpoints and on chat. The token had already been saved, but the step that writes a default model had not run, and a home with no default model cannot run a turn. Login now waits for both, and retries the provisioning step up to twice, five seconds apart, without opening the browser again.

The same day found a quieter bug. Luminair had been looking for the wrong token file. The CLI writes its login to credentials/kimi-code.json; files ending in -tokens.json belong to its MCP sign-ins. So every successful Kimi login had been reported as failed.

06A meter from the vendor's source

Read the CLI, not the docs.

Every Luminair account shows how much of its plan is left. For Kimi, rather than guess at an API, we looked at the CLI: it has its own usage panel, and it ships as readable JavaScript. On 15 September the engine's meter was written by reading the bundled 0.36 source for the call that panel makes.

Fig 03  How the Kimi meter is filleddesktop/engines/kimi.js
ACCOUNT HOME credentials/ kimi-code.json token API.KIMI.COM GET /coding/v1/usages GET /coding/v1/me ACCOUNT ROW plan name, email METER ROWS Weeksummary: used / limit, reset time 5ha rolling window from the limits list Boosterwallet balance, fixed-point ×1e6 Bars left empty on purpose: no Kimi account was signed in when the meter was written. A 401 reads “Kimi sign-in expired; the next Kimi turn refreshes it”: a soft failure.
Schematic of the code path. The three row names and their order are the ones the test suite asserts against a stubbed server. Luminair reads the CLI's token but never refreshes it; the CLI does that on its next turn.

The code is candid about its limits. A warning in the file says it was “Written from the CLI's source; no Kimi account was signed in on this Mac to see a live payload”, and tells the next person to check it on the first real sign-in. The test suite pins the shape: a stubbed /usages and /me must become the rows Week, 5h and Booster, and fill in the plan name. For a vendor with no usage API at all, see A usage meter for a vendor that has none.

07A CLI that moves faster than you

0.36 to 2.1.1 in a week.

Bundling a vendor's CLI buys you their login, their thread format and their tools. It also ties you to their release pace, and Kimi's moved quickly. Dependabot watches Luminair's dependencies and opens a pull request when a newer version appears. Here is what it proposed for @moonshot-ai/kimi-code:

Fig 04  Proposed Kimi Code upgradesReal data · Dependabot PRs
6 SEP 13 SEP 27 SEP bundled: 0.36.0, all month 0.41.0 PR 15 superseded 2.1.1 PR 23, one week later superseded 2.1.1 PR 32 closed, release ignored
From the repository's pull request history, September 2026. Hollow dots are proposals that were never merged. The 2.1.1 jump is a major-version change, as Dependabot itself labelled it.

A jump from 0.x to 2.x is exactly the kind of change that can move a token file, rename a flag or change a stream format, which are the three things the Kimi lane depends on. So the app stays on the version its notes were verified against, and it sets KIMI_CODE_NO_AUTO_UPDATE=1 on every call so the CLI never replaces itself under a running app. Moving to a new version means re-checking the login, the token path and the meter, not just bumping a number.

The trade-off, plainlyPinning keeps the lane working. It also means Luminair's Kimi runs on an older CLI than Moonshot's latest, and any fixes in newer releases wait until someone does the re-verification.
08Find it in the app

Three clicks to a new model.

  1. 1For the local models, open your avatar at the bottom left, then Settings › Models. Open Qwen, GLM or DeepSeek and press its download button, for example Download DeepSeek R1. Each is about 18 to 20 GB.
  2. 2For Kimi, add a Kimi account from the account menu and finish the device sign-in in your browser. You need a Kimi Code plan.
  3. 3Pick the model from the session's model menu. You can switch to it mid-thread, and back.

One limit to know: Kimi sessions run on your computer, not from the phone. The app's relay tells you so and suggests sending the turn from the desktop.

09What we checked

Checked, and not claimed.

6/6
Engine usage-meter tests pass, including Kimi's Week, 5h and Booster rows
0.36.0
The Kimi Code version bundled and verified; three upgrade PRs, none merged
3
Chinese model families that run entirely on your computer, with no key
2
Provisioning retries after a Kimi server error, five seconds apart

What this post does not claim

  • How these models compare with each other or with US models. We have not benchmarked them, and the scores quoted above are the sources' own.
  • That the Kimi meter has been seen against a live account. Its code says it has not, as of when it was written.
  • Which model a Kimi Code plan runs. Luminair sends no model flag and uses the plan's default.

Sources

Try a model from another lab

Download Qwen, GLM or DeepSeek, or sign in with Kimi Code, and switch mid-thread.

Download Luminair →