The short version
- “Vibe coding” means letting the model write code you do not read. It is great for prototypes. Andrej Karpathy, who coined it, now says serious work needs agentic engineering: agents do the typing, people keep the oversight.
- Luminair writes that oversight into every turn as abilities: Plan first, Surgical edits and Proof of done. All three are on by default and bind on every engine.
- Some rules are enforced, not just asked for. On the Claude lanes, a whole-file rewrite that is really a small patch is refused, and every edit lists the files that import the one you changed.
- Our own release pipeline follows the same idea: tests gate the build, every file is checked by sha256 in a hidden draft, and only then does one step make it public.
The year of the vibes.
In February 2025 Andrej Karpathy gave a name to something many people were already doing. As quoted in full on Simon Willison's blog: “There’s a new kind of coding I call “vibe coding”, where you fully give in to the vibes, embrace exponentials, and forget that the code even exists.” And: “I “Accept All” always, I don’t read the diffs anymore.”
It caught on because it works, for some things. A weekend tool, a one-off script, a prototype you will throw away. Willison, in the same post, drew the line that most working developers recognise: “My golden rule for production-quality AI-assisted programming is that I won’t commit any code to my repository if I couldn’t explain exactly what it does to somebody else.”
The data from that summer was a cold shower. In a randomized controlled trial with experienced open-source developers working on their own repositories, METR found that “When developers are allowed to use AI tools, they take 19% longer to complete issues”, and that “developers expected AI to speed them up by 24%, and even after experiencing the slowdown, they still believed AI had sped them up by 20%.” The feeling of speed and the fact of speed had come apart.
Raise the ceiling.
A year on, Karpathy offered a second name. His own summary of a talk at Sequoia Ascent 2026 puts the two side by side: “Vibe coding raises the floor.” And: “Agentic engineering raises the ceiling. It is the professional discipline of coordinating fallible agents while preserving correctness, security, taste, and maintainability.” His verdict is practical rather than preachy: “Vibe coding is fine for prototypes and personal tools. Agentic engineering is what serious teams need.”
We agree, with one addition. Discipline that lives only in a person's head does not survive a long day, a tired evening or a model that says “done” with confidence. So we wrote it down where every model reads it on every turn, and where possible we made the app enforce it.
In Luminair these rules are called abilities. They are short blocks of instruction Luminair adds to the session contract, stored on the Mac so they apply on every lane: the desktop, the phone relay and every command line engine. You can switch each one off. They start on.
Solace
Say it back before you build it.
The most expensive bug is building the wrong thing well. Plan first targets that, and only that. It applies to “a request that builds a new feature, changes behaviour across more than one file, or leaves open what "done" means.” A quick fix, a one-file change or a question skips it, so small work never pays for a planning step.
For a qualifying request the model must, before any code: say in one or two sentences what it understood; ask at most three questions, and only ones that truly block the work; then give a short numbered plan in which each step names the file it touches “and how you will check that it worked.”
One line matters as much as the rest: “The plan is information, not a permission request: never ask the user to approve it”. Plan first is meant to make intent visible, not to add a sign-off loop. If the work drifts from the plan, the model has to say so when it happens, not at the end.
Patch, don't rewrite.
Ask a model to change three lines in a 500-line file and it will sometimes send back all 500. That costs, in the ability's words, “3 to 5 times the tokens of the same change as a patch and risks silently dropping code.” The ability tells the model to use exact edits, batch changes to one file in one pass, read an unfamiliar file’s skeleton first and open only the span it needs, and never re-read a file it just edited.
On the Claude lanes, Luminair does not just ask. A guard sits in front of the Write tool and compares what the model wants to write with what is already on disk.
The second half comes from a well-known failure pattern: an agent fixes one file and misses the files that depend on it. Right after an edit lands, Luminair runs a bounded git grep for files that import the changed file and hands the list to the model, with an instruction to act on it. It skips generic names like index or utils that would match everything. The Claude command line lane gets the same guard and list through hook scripts; other engines get the rule as text.
“Done” is a claim. Prove it.
The last ability targets the word models say too easily. Proof of done begins: “never call a change done, fixed, working, or shipped on the strength of having written the code.” Before the claim, the model runs the check that proves it: the tests for that area, a build, a direct run, or a read-back of the saved result. A behaviour change in a project with tests gets a test “that fails without the change and passes with it.”
Two details make it useful rather than ceremonial. Every new control has to be traced end to end: the element, the handler it fires, the state it writes, the code that reads it, because “A control wired to nothing is not done.” And the parts that could not be checked, because they need a phone or a signed-in app, are named: “An unverified part is reported as unverified, never as done.”
Solace adds gates on top for tracked work. With done-enforcement on, an item marked tested needs real evidence, and only you can mark an item confirmed. An optional independent verifier starts a fresh model with no prior context to trace a finished feature through the code.
Nothing is public until it is verified.
The same principle runs our releases. On 3 September 2026, after what the workflow file calls “a day lost to partial releases”, we replaced several scripts with one pipeline whose header states its promise: “nothing becomes visible until every artifact exists and verifies.”
A few of its rules, each written into the workflow. The version bump and tag are made from a fresh checkout, so “a laptop's dirty working tree” can never leak into a release. Each Mac build is checked after packaging, including whether the bundled engine really spawns. Every file lands in a draft that nobody outside can see, and the public “latest” pointer only moves after every expected asset is verified “by name, size and sha256.” After the one publish step, the workflow downloads the first byte of every file the website's buttons point at, plus both update feeds, and fails loudly if any is missing: “a site button is DEAD”.
It is the same shape as Proof of done, applied to shipping. Writing the code is not the finish line. The check is.
Checked, and not claimed.
What this post does not claim
- That these abilities make any model faster or its code better by a measured amount. We have not run that study.
- That an ability is a guarantee. Only the Write guard and the release gates are enforced by code; the rest are instructions a model follows well or badly.
- That vibe coding is wrong. For prototypes and personal tools it is often the right call, and every switch here can be turned off.
Sources
- Simon Willison's Weblog · 19 March 2025Not all AI-assisted programming is vibe coding (but vibe coding rocks)
- METR · 10 July 2025Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
- Andrej Karpathy · 30 April 2026Sequoia Ascent 2026 summary
For how Luminair runs several agents at once without losing track, read Many agents, one Mac.
Keep the speed, add the proof
Plan first, patch don't rewrite, and never call it done without a check.