← Blog Practice 25 September 2026 9 min read

Vibe coding vs agentic engineering.

In 2025 we learned to let the model write the code. In 2026 the question became who checks it. Here is how Luminair turns the discipline part into switches and gates you can see: plan first, patch don't rewrite, prove it before you call it done, and never publish what has not been verified.

Fig 01  Where the discipline sits, request to releaseSchematic · from the code
YOU “Add an export.” 1 · BEFORE CODE Plan first Bigger requests only What I understood Up to 3 blocking questions Numbered plan, each step with its check Never a permission request 2 · WHILE CODING Surgical edits Patch, never re-emit a file Guard refuses a rewrite that is really a patch After an edit: the files that import it are listed Enforced on the Claude lanes 3 · BEFORE “DONE” Proof of done Run the check first A test that fails without the change Proof next to the claim Unverified says unverified 4 · TRACKED Done gates “Tested” needs real evidence A fresh verifier traces the feature Only you confirm 5 · SHIPPING Release Tests gate Build, sign Draft sha256 every file Publish once Check the site Steps 1 to 3 are session abilities, on by default, sent to every engine. Step 4 is Solace. Step 5 is Luminair's own release pipeline.
Drawn from lib/mc-behaviours.js, the Write guard in desktop/main.js, the Solace features in desktop/renderer.js and .github/workflows/release.yml. Labels are shortened; the wording each model receives is quoted in the post.

The short version

  1. “Vibe coding” means letting the model write code you do not read. It is great for prototypes. Andrej Karpathy, who coined it, now says serious work needs agentic engineering: agents do the typing, people keep the oversight.
  2. Luminair writes that oversight into every turn as abilities: Plan first, Surgical edits and Proof of done. All three are on by default and bind on every engine.
  3. Some rules are enforced, not just asked for. On the Claude lanes, a whole-file rewrite that is really a small patch is refused, and every edit lists the files that import the one you changed.
  4. Our own release pipeline follows the same idea: tests gate the build, every file is checked by sha256 in a hidden draft, and only then does one step make it public.
02Forget the code exists

The year of the vibes.

In February 2025 Andrej Karpathy gave a name to something many people were already doing. As quoted in full on Simon Willison's blog: “There’s a new kind of coding I call “vibe coding”, where you fully give in to the vibes, embrace exponentials, and forget that the code even exists.” And: “I “Accept All” always, I don’t read the diffs anymore.”

It caught on because it works, for some things. A weekend tool, a one-off script, a prototype you will throw away. Willison, in the same post, drew the line that most working developers recognise: “My golden rule for production-quality AI-assisted programming is that I won’t commit any code to my repository if I couldn’t explain exactly what it does to somebody else.”

The data from that summer was a cold shower. In a randomized controlled trial with experienced open-source developers working on their own repositories, METR found that “When developers are allowed to use AI tools, they take 19% longer to complete issues”, and that “developers expected AI to speed them up by 24%, and even after experiencing the slowdown, they still believed AI had sped them up by 20%.” The feeling of speed and the fact of speed had come apart.

The feeling of speed and the fact of speed had come apart.
03Engineering, again

Raise the ceiling.

A year on, Karpathy offered a second name. His own summary of a talk at Sequoia Ascent 2026 puts the two side by side: “Vibe coding raises the floor.” And: “Agentic engineering raises the ceiling. It is the professional discipline of coordinating fallible agents while preserving correctness, security, taste, and maintainability.” His verdict is practical rather than preachy: “Vibe coding is fine for prototypes and personal tools. Agentic engineering is what serious teams need.”

We agree, with one addition. Discipline that lives only in a person's head does not survive a long day, a tired evening or a model that says “done” with confidence. So we wrote it down where every model reads it on every turn, and where possible we made the app enforce it.

In Luminair these rules are called abilities. They are short blocks of instruction Luminair adds to the session contract, stored on the Mac so they apply on every lane: the desktop, the phone relay and every command line engine. You can switch each one off. They start on.

Solace · Controls
Solace
RequestsControlsSpend
Reply style
Surgical editsSaves tokensFiles change by targeted patch, never by re-emitting the whole file. On Claude, patch-like whole-file writes are refused and supported edits trigger an import search.
Work discipline
Plan firstFreeFor a bigger request: says what it understood, asks only what blocks it, shows a short plan, then builds.
Proof of doneFreeNothing is called done until a check was run. The proof is reported next to the claim.
Drawn with the row names, tags and one-line descriptions from renderMcControlsTab in desktop/renderer.js; other rows in the tab (Lean replies and the quality gates) are left out. Solace is part of Pro+.
04Plan first

Say it back before you build it.

The most expensive bug is building the wrong thing well. Plan first targets that, and only that. It applies to “a request that builds a new feature, changes behaviour across more than one file, or leaves open what "done" means.” A quick fix, a one-file change or a question skips it, so small work never pays for a planning step.

For a qualifying request the model must, before any code: say in one or two sentences what it understood; ask at most three questions, and only ones that truly block the work; then give a short numbered plan in which each step names the file it touches “and how you will check that it worked.”

One line matters as much as the rest: “The plan is information, not a permission request: never ask the user to approve it”. Plan first is meant to make intent visible, not to add a sign-off loop. If the work drifts from the plan, the model has to say so when it happens, not at the end.

05Surgical edits

Patch, don't rewrite.

Ask a model to change three lines in a 500-line file and it will sometimes send back all 500. That costs, in the ability's words, “3 to 5 times the tokens of the same change as a patch and risks silently dropping code.” The ability tells the model to use exact edits, batch changes to one file in one pass, read an unfamiliar file’s skeleton first and open only the span it needs, and never re-read a file it just edited.

On the Claude lanes, Luminair does not just ask. A guard sits in front of the Write tool and compares what the model wants to write with what is already on disk.

Fig 02  When a whole-file write is refusedThresholds from surgicalWriteVerdict
Refused: a patch posing as a rewrite redo it as Edit calls Passes: mostly new content under 30 lines: always passes 100% 60% 40% 0% 0 30 400 lines in the file → Vertical: share of the existing lines kept verbatim. New files always pass.
The rule is real; the horizontal axis is not to scale. Lines are counted after trimming, ignoring lines of two characters or fewer. A refusal tells the model why and what to do instead, and the panel counts refusals per day.

The second half comes from a well-known failure pattern: an agent fixes one file and misses the files that depend on it. Right after an edit lands, Luminair runs a bounded git grep for files that import the changed file and hands the list to the model, with an instruction to act on it. It skips generic names like index or utils that would match everything. The Claude command line lane gets the same guard and list through hook scripts; other engines get the rule as text.

06Proof of done

“Done” is a claim. Prove it.

The last ability targets the word models say too easily. Proof of done begins: “never call a change done, fixed, working, or shipped on the strength of having written the code.” Before the claim, the model runs the check that proves it: the tests for that area, a build, a direct run, or a read-back of the saved result. A behaviour change in a project with tests gets a test “that fails without the change and passes with it.”

Two details make it useful rather than ceremonial. Every new control has to be traced end to end: the element, the handler it fires, the state it writes, the code that reads it, because “A control wired to nothing is not done.” And the parts that could not be checked, because they need a phone or a signed-in app, are named: “An unverified part is reported as unverified, never as done.”

Solace adds gates on top for tracked work. With done-enforcement on, an item marked tested needs real evidence, and only you can mark an item confirmed. An optional independent verifier starts a fresh model with no prior context to trace a finished feature through the code.

Tests are the gateLuminair's own desktop app has 190 core test files. Every push runs them on a CI gate, and a release cannot start building until they pass serially, with a security audit and a software bill of materials, on each Mac build lane.
07A release that checks itself

Nothing is public until it is verified.

The same principle runs our releases. On 3 September 2026, after what the workflow file calls “a day lost to partial releases”, we replaced several scripts with one pipeline whose header states its promise: “nothing becomes visible until every artifact exists and verifies.”

Fig 03  One release, end to endJobs from release.yml
cut secrets checked, clean tag Mac arm64 · tests, build Mac Intel · tests, build Windows · tests, build Linux · best effort draft (hidden) finalize: name, size, sha256 publish one step site every button A missing credential fails in cut, about a minute in, before any 30-minute build starts.
Simplified from the job graph drawn at the top of .github/workflows/release.yml. Signing, notarisation and the single-arch carry-forward paths are left out.

A few of its rules, each written into the workflow. The version bump and tag are made from a fresh checkout, so “a laptop's dirty working tree” can never leak into a release. Each Mac build is checked after packaging, including whether the bundled engine really spawns. Every file lands in a draft that nobody outside can see, and the public “latest” pointer only moves after every expected asset is verified “by name, size and sha256.” After the one publish step, the workflow downloads the first byte of every file the website's buttons point at, plus both update feeds, and fails loudly if any is missing: “a site button is DEAD”.

It is the same shape as Proof of done, applied to shipping. Writing the code is not the finish line. The check is.

08What we checked

Checked, and not claimed.

3
Discipline abilities, on by default: Plan first, Surgical edits, Proof of done
7/7
Solace behaviour tests pass, including the Write guard
190
Core test files that gate every release build
60%
Share of a 30-line-plus file kept verbatim at which a whole-file write is refused

What this post does not claim

  • That these abilities make any model faster or its code better by a measured amount. We have not run that study.
  • That an ability is a guarantee. Only the Write guard and the release gates are enforced by code; the rest are instructions a model follows well or badly.
  • That vibe coding is wrong. For prototypes and personal tools it is often the right call, and every switch here can be turned off.

Sources

For how Luminair runs several agents at once without losing track, read Many agents, one Mac.

Keep the speed, add the proof

Plan first, patch don't rewrite, and never call it done without a check.

Download Luminair →