H
Howardism
Plate IIAI Coding Practice中文HOWARDISM

Vibe Coding vs. Agentic Engineering

Vibe coding raises the floor (anyone builds); agentic engineering preserves the quality bar while going faster; ">10x and widening"; hire on big projects, not puzzles

Article metadata
Publication details
Published:May 23, 2026
Filed:Concept
Domain:AI Coding Practice
Reading:15 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Vibe Coding vs. Agentic Engineering

Sources#

Summary#

Andrej Karpathy coined "vibe coding" in 2025 and, a year later, names its serious successor: agentic engineering. The distinction is about which bar moves. Vibe coding raises the floor — anyone can build software now. Agentic engineering preserves the quality bar of professional software while going much faster: "you're not allowed to introduce vulnerabilities due to vibe coding; you're still responsible for your software, but can you go faster — and how do you do that properly?" It's an engineering discipline for coordinating spiky, fallible, stochastic-but-powerful agents without sacrificing quality.

The two bars#

  • Vibe coding — floor up. Everyone can vibe-code anything. "Amazing, incredible." Democratization (cf. Printing Press Software Democratization). Quality is not the point; access is.
  • Agentic engineering — ceiling up, quality held. You keep the responsibilities of professional software (security, correctness, maintainability) and use agents to go faster without dropping below that bar. "Doing that well and correctly is the realm of agentic engineering."

These are different activities, not points on one line. One lowers the entry cost; the other raises the output ceiling for people who already clear the bar.

"10x is not the speedup"#

Karpathy explicitly retires the old "10x engineer" trope as too small: "10x is not the speedup you gain… people who are very good at this peak a lot more than 10x." The ceiling on agentic-engineering capability is very high, and the spread between mediocre and AI-native practitioners widens, not narrows. (Echoes Harness Shrinkage as Models Improve: the leverage keeps growing as models improve; the binding constraint becomes the operator's taste — see Outsource Your Thinking, Not Your Understanding.)

What an AI-native practitioner looks like#

Asked to contrast a mediocre vs. a fully AI-native user of cloud code / codex / open claw, Karpathy's answer is mundane and important: invest in your setup, use all the tool's features. Same as the engineers who got the most out of Vim or VS Code — now applied to Claude Code / Codex. Mastery is configuration-and-features fluency, not a secret prompt.

Hiring has to be refactored#

A practical corollary: most teams still hire with the old paradigm (puzzles, leetcode). Karpathy argues agentic-engineering hiring should look like "give me a really big project and watch someone implement it well" — e.g., build a secure Twitter-clone-for-agents, then a red-team agent ("codex 5.4 xhigh") tries to break it and can't. Hiring should test verifiable, end-to-end build-and-defend ability, not isolated puzzle-solving. (See The Verifiability Thesis for why "and it can't be broken" is the load-bearing half.)

The human residue#

Even at the high ceiling, the human stays in charge of spec, taste, judgment, and oversight — agents do the fill-in-the-blanks. His MenuGen war story: the agent matched Stripe and Google accounts by email address instead of a persistent user ID — "such a weird thing to do," the kind of mistake Jagged Intelligence (Ghosts, Not Animals) predicts. You must design the spec ("these must be unique user IDs we tie everything to") and supply the taste; the agent handles the API details you've stopped memorizing.

The veteran's failure mode: over-specification (Cherny)#

Boris Cherny names the characteristic mistake of experienced engineers with modern models (YC interview, July 2026, practitioner-opinion): "way over-specific instructions… you must do one, then two, then three, then four. And for modern models, that's actually really not the way to do it." The correction: go a level higher — describe the task, the guardrails, and the exit criteria, then "let the model cook." He frames it as an unlearning problem: decades of building deterministic systems trained engineers to over-specify, and "when I look at engineers that have been coding for decades, this is a really really common failure mode… it's a journey to unlearn it" — treat the model "like you would a coworker. That's the level of intelligence that it's at now." This is the same operator-skill axis as Karpathy's mediocre-vs-AI-native spread, but with the sign flipped: what widens the gap isn't tool fluency alone, it's willingness to drop priors that used to be professionalism. (The advice is model-generation-indexed: "this is just not something that would have worked 6 months ago, but it does work today" — the delegation level itself is a shrinking-harness dial.)

The moving goalposts: supervised vs. unsupervised (Ambrosino)#

Andrew Ambrosino (OpenAI Codex) restates the same "which bar moves" distinction as a moving-goalposts observation, and welcomes the movement as evidence of progress. Asked what fraction of the product is AI-written: "If you use the goalposts from last year, 100% of our product is AI-written code. So the question is more like — is the code written supervised versus unsupervised? That's a totally different thing. I welcome the moving of goalposts, because that means we're making progress." The salient axis is no longer human-vs-AI authorship (settled) but how much human supervision the authoring still needs — which is exactly the agentic-engineering "quality bar held while going faster" line, measured as supervision cost.

He also captures the interactive form as "coding is steering the AI": the honest measure of AI's contribution isn't "what percent of my code did AI write" but "how many times did I have to steer it in the right direction" — the allocator/steerer role, restated. And on the frontier: "loops are so last week." The leading edge has moved past orchestrated loops to autonomous development and harness engineering — e.g. an agent doing overnight "garbage collection" of the codebase — though he flags it isn't there yet (models "usually increase complexity" and are bad at deleting code). This dates the vibe-coding→agentic-engineering ladder from the OpenAI side: the practitioner question is now supervised-vs-unsupervised and how autonomous a loop you can trust, not can the model write it.

The dissent from the craft side: DHH rejects the vocabulary and redraws the line#

DHH — twenty years of handcrafted Ruby, now shipping 100% agent-written code — engages this page's framing directly on Lex Fridman #501 (2026-08-26, practitioner-opinion) and rejects half of it.

On the term. "Agentic engineering — oh, I fucking hate that term… first of all, it's become marketing slop speak at this point. Like, it's just slapped onto everything. I wish we had a different word that just meant AI doing stuff." He is no happier with the other one: vibe coding "smells exactly like script kiddies did in the early 2000s. People applying PHP scripts they just downloaded offline that they don't understand anything of."

On the line itself — a different cut. Karpathy's split is about which bar moves (floor vs. ceiling). DHH's is about implementation visibility, and it is sharper because it is binary: "vibe coding, if we define it here, is you tell an agent to build software for you. You do not look at the implementation. That, to me, is what separates vibe coding from programming or, let's say, agent-accelerated development." He refuses to let the result be called programming, on the grounds that programming names an understanding of primitives — loops, conditions, variables — not the production of programs: "you could say in the pre-agentic era, well, your CEO is programming… I don't think most people would call that CEO a programmer." He then classifies his own work under both labels at once: he vibe-coded Omawrite (a C++/Qt Markdown editor he has never read a line of, deliberately, "100% a black box") while still reviewing the shape of everything in Omarchy's model layer.

On expertise as a liability. The page's implicit assumption — that the AI-native practitioner is the strong engineer who invested in setup — takes its most direct hit here. Asked whether there are problems where programmers do worse than non-programmers at this, DHH says "100%." His reasoning: "there's a lot of programmers who are not very good product managers. Software is product management. What should it do? Who should it do it for?… And in the agentic era where you are letting an agent do the implementation, these are the skills you need." He reports it as autobiography, not theory: "I actually think for a while it was to my deficit to know as much as I know about programming, because I was instructing the agents to do things as I prescribed them to do… the next moment allowed anyone to describe outcomes and get better solutions than if you had a programmer prescribe the path." Set this against Returns to Expertise in Agentic Coding, which measures the opposite sign — domain expertise doubles verified success — and the reconciliation is that DHH is describing implementation expertise misapplied as prescription, while the Anthropic study measures domain expertise applied as context.

Corroborating Cherny's over-specification finding, from outside Anthropic. DHH cites the 80% system-prompt cut by name — "one of the things that Boris, working on Claude Code, shared… was that the system prompt that they ship for Opus 5 shrunk by 80% because the agent not only needed far less human instruction, it was actually being damaged by overly prescriptive humans" — and supplies the analogy that makes it stick: "any programmer who's had a pointy-haired boss knows exactly what that is like… You write shittier code if you're mandated to do things that are against your better judgment. Why would an agent not be the same?" See Harness Shrinkage as Models Improve.

And an older lineage for it. His positive prescription is not "prompt better," it is the agile argument transplanted: specification-up-front failed for forty years because "no one knows what they want until they receive it. You don't know what a program should do until you play with it. So in the agentic age, you should resist the temptation to be overly specific upfront. Be as vague as you can to manifest something, then interact with the something." The human contribution he names is differential evaluation — give a person three options and the gut answers instantly; give them twenty-two and they fall off a cliff. That is a concrete shape for the "taste" the rest of this page treats as a primitive.

Connections#

Open Questions#

  • Karpathy hints at "one domain that's very [valuable]" for founders but won't say which (didn't want to "vague-post on stage"). What verifiable RL-environment domain is he gesturing at?
  • If the mediocre/AI-native spread keeps widening, what does that do to team composition — a few extreme outliers plus agents, vs. broad mid-level staffing?

Sources#

§ end
Cited by 28
Related articles
  • Verification as the New Bottleneck

    Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…

  • Agentic Technical Debt

    Debt that *compounds* (not just accumulates) because each agentic-coding session re-derives architectural decisions wit…

  • Claude Code

    Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…

  • Harness Shrinkage as Models Improve

    Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…

  • Andrej Karpathy

    Co-founder OpenAI, ex-Tesla AI, Eureka Labs; coined "vibe coding," Software 1/2/3.0, "ghosts not animals," "agentic eng…