Sources#
Summary#
Jeff Dean's answer (YC Startup School 2026, practitioner-opinion) to the question every founder building on frontier models has to answer — will the model eat this? — is an empirical test you can run tonight rather than a forecast:
"Look at what the current more general models can do in that problem domain. You can test them with, are they able to do this thing very well? And if they're completely failing, that's probably a good sign. If they're kind of able to do some of it but not very well, that's maybe not a great sign, because that's probably a sign that the capability is starting to be present in those models… So look for something where the model succeeds 0% or 1% of the time, not 20%."
The rule inverts the naive read of a demo. Partial success is the danger signal, not the encouraging one. A 20% success rate means the capability exists in latent form and needs only more data, more scale, or more inference budget — all of which are arriving on someone else's roadmap. A 0% success rate means something structural is missing, and structural gaps are the only ones with a chance of outlasting two release cycles.
The paired question Dean attaches to it is a horizon, not a binary: "do you think the models at the forefront are going to get better at that in the next six months or 12 months? Or is it something they're not going to be able to do for a couple of years or three years?" The rule is a cheap instrument for estimating that horizon, because the current success rate is a proxy for how far the capability is from the frontier's leading edge.
What survives the test#
Dean names two shapes that plausibly clear it, and both are shapes where the general model is missing an input rather than an ability:
- Data the general model structurally cannot see. His example is a product that organizes a user's own personal information: "the model won't necessarily have access to that." The advantage isn't cleverness, it's visibility. This is Compounding Data Moat's argument arriving from the model side — the moat is a corpus the frontier lab's training set does not and will not contain.
- A niche model trained on data nobody else has. AlphaFold is his archetype: "a very specific model for protein folding… not a general model," and a durable one. He extends the shape to materials science and chip design. The economics he claims for it are the interesting part — "maybe it doesn't take that much compute to train a niche model for this particular problem, but you can get something that's highly accurate." Note this is a
practitioner-opinionclaim from someone with Google-scale compute intuitions; the corpus has no measurement of what a defensible niche model actually costs a startup in 2026.
He caps both with the honest caveat: "the general models are definitely getting better at a broader and broader range of things." And he puts a criterion above the test — pick something you're excited about and that would be useful — explicitly ranked as "the number one selection criteria," with the capability test second.
The collision with build-for-the-next-model#
This page and Build for the Next Model read the same signal — the model almost does it — and return opposite verdicts.
| Signal | Verdict | |
|---|---|---|
| Build for the Next Model (Carey, Cherny, Cat Wu, Ambrosino) | The model gets this ~20% right | Good. Prototype it; the next release closes the gap for free |
| This page (Dean) | The model gets this ~20% right | Bad. The capability is arriving; don't found a company on it |
The contradiction is real and neither source addresses the other. The reconciliation that survives both is who owns the surface the release lifts:
- Build-for-the-next-model is about a capability gap inside a product you already own. Claude Design's unsolved problems were fixed by Opus 4.7 landing under a product Anthropic controlled — the tide lifted their boat, and they captured the gain.
- The 1% rule is about what the company is for. If the general model closes your gap and the surface is generic, the tide lifted the frontier lab's boat and floated it into your market. The release is a subsidy in the first case and a competitor in the second.
Stated as one rule: wait for the model on features, never on the wedge. Both sources agree the release cadence is forecastable; they disagree only on whether you are standing on it or in front of it. The same asymmetry is what Harness Shrinkage as Models Improve describes at the scaffolding layer — capability migrating inward is free if you own the thing it migrates into.
There is a second, softer tension with Problem-Solution Fit Discipline. That page's discipline is to validate the problem before building; the 1% rule validates durability and says nothing about whether anyone wants the thing. A 0%-success problem nobody has is still a bad business. The two tests are orthogonal and both are required — run the market test first, since Dean himself ranks passion and usefulness above the capability screen.
Where the rule is weak#
- It is a snapshot of a curve, taken once. A 0% rate today can be 40% after one release if the gap turns out to be elicitation rather than capability — which is exactly what Latent Capability Overhang says is routinely true of already-released models. The honest version of the test measures the trend across two or three generations, not the level at one moment. Dean's own framing (six months / twelve months / three years) implies a trend, but the instrument he gives measures a point.
- 0% can mean "impossible," not "defensible." The rule cannot distinguish a task the frontier will not reach for three years from one nobody can do at all. It has no notion of whether the problem is solvable by the founder either.
- It is unmeasured. One practitioner's heuristic, stated once, with two supporting examples chosen after the fact (AlphaFold succeeded; the failures aren't enumerated). Nothing in the corpus tests whether 20%-success markets actually get absorbed faster than 0%-success ones.
Connections#
- Jeff Dean — the source; the rule sits inside a broader method of first-principles bottleneck estimation
- Build for the Next Model — the direct contradiction, reconciled above on whether you own the surface the release lifts
- Narrow Wedge into a Legacy Market — the complementary axis: that page picks a wedge against an incumbent product's unused surface, this one picks against the frontier model's capability curve. A wedge has to clear both
- Compounding Data Moat — the durable version of "data the general model cannot see"; the 1% rule is a screen for it, the moat page is what you do once through the screen
- Problem-Solution Fit Discipline — the orthogonal test: durability against models says nothing about demand
- Latent Capability Overhang — the rule's main failure mode: a 0% reading can be unelicited capability rather than absent capability
- Harness Shrinkage as Models Improve — the same inward-migration dynamic one layer down, where it is a gift rather than a threat because you own the harness
Open Questions#
- Does the rule hold empirically? Nothing here tests whether markets where models scored ~20% in 2025 were absorbed faster than markets where they scored ~0%. The data to check it (benchmark-era capability snapshots against startup outcomes) exists in principle.
- What is the 2026 cost of a defensible niche model? Dean asserts "maybe it doesn't take that much compute"; a founder needs the number, and the corpus doesn't have it.
Sources#
- Jeff Dean: The 1% Rule for Building in AI — YC Startup School 2026, Jeff Dean with Diana Hu (2026-07-30,
practitioner-opinion): §"Where Startups Can Still Beat Google" — the 0%-or-1%-not-20% test, the six/twelve-month/three-year durability horizon, the personal-data and AlphaFold-shaped exceptions. COI: Google's Chief Scientist advising founders on which markets Google will not enter
Cited by 8
- Build for the Next Model×2
A sharper tension, from outside Anthropic: Jeff Dean tells founders to read "the model almost does…
- Jeff Dean×2
His most quotable contribution here, and the one with its own page: test the general models on your…
- Compounding Data Moat
One Percent Rule Wedge Selection — the screen that selects for this moat before you build: Dean's…
- Latent Capability Overhang
One Percent Rule Wedge Selection — the overhang is that rule's main failure mode: a founder reading…
- Startup & Founder
One Percent Rule Wedge Selection — Jeff Dean's test for what a startup should build: run the…
- Narrow Wedge into a Legacy Market
One Percent Rule Wedge Selection — the other axis a wedge has to clear: this page picks against an…
- Open Questions Backlog
One Percent Rule Wedge Selection ×2 (oldest 3d) — Does the rule hold empirically?
- Problem-Solution Fit Discipline
One Percent Rule Wedge Selection — the orthogonal screen: this page tests whether the problem is…
Related articles
- AI-Native Startup Lifecycle
Anthropic's May 2026 reframing of Idea/MVP/Launch/Scale assuming AI infrastructure: each stage's headcount/capital/skil…
- Build for the Next Model
Prototype the thing that almost works, not the thing that already works: bet that the next concrete model release (not…
- Claude Code
Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…
- Founder as Agent Orchestrator
Founder role shift: less individual contributor, more orchestrator of specialized AI assistants; non-technical founders…
- Narrow Wedge into a Legacy Market
Disrupt without being feature-complete: be the best for a narrow customer profile (tech cos outgrowing QuickBooks); Goo…
