H
Howardism
Plate IIEvals & Benchmarks中文HOWARDISM

Evaluation Horizon Versus Release Cadence

Noam Brown's observation that the horizon a frontier model can operate over is growing faster than the interval between frontier releases, so there will be no window in which a model can be safety-evaluated at the full length of its own capability before its successor ships — a structural expiry on pre-release evaluation that the corpus's own budget and time-horizon curves already point at, plus the flip side Brown says has no good answer: slowing the cadence to buy evaluation time widens the gap between what a lab holds internally and what the world can use

Article metadata
Publication details
Published:September 20, 2026
Filed:Concept
Domain:Evals & Benchmarks
Reading:18 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Evaluation Horizon Versus Release Cadence

Sources#

Summary#

Every pre-release safety evaluation carries an unstated assumption: that you can evaluate the model in less time than you have before you need to ship it. Noam Brown (OpenAI, Dwarkesh Podcast, 2026-09-17, practitioner-opinion) argues that two trend lines are converging on the point where that stops being true, and that almost nobody is planning for it:

"You want it to do a week-long task, it can do a week-long task. We'll probably get to the point where they can do month-long tasks. We'll probably get to the point where they can do 3-month-long tasks. If you're in a world where they can operate effectively over three months, but the model release cycle is every two months, then you don't have a way to evaluate the models at the full length of their capabilities before the next model release cycle."

This is Brown's claim, not a measurement — he names no evaluation duration, publishes no curve, and says explicitly that "this isn't an issue right now, but it is quickly becoming an issue that we have to figure out a solution for." Treat the arrival date as a forecast. What makes it worth a page rather than a line is that the corpus already measures both trend lines separately and has never put them on the same axis.

The two lines#

Cadence. Brown's figure for the release interval is "at most every two months, sometimes faster," with "every week there's a new AI breakthrough." This is the one quantity an outsider can check, and the wiki's model pages are the check.

Horizon. METR's time-horizon curve is the quantitative version of Brown's week → month → three-month ladder, and it is the load-bearing external corroboration: reliable task length doubling roughly every four months is a faster exponential than the release interval is shrinking. Brown's own framing of the same growth is the flat contrast with GPT-3 — "you could loop it to do stuff over long horizons. You just wouldn't do very well at it" — so what changed is not that long runs became possible but that they became effective, which is exactly what makes them require evaluation at length.

The crossing is not a single date, because the two quantities are not the same kind of thing. A model's horizon is a distribution over tasks, not a scalar; a release cadence is an operational choice. What Brown is pointing at is the ordering, and the ordering is already visible in the corpus: Large-Scale Test-Time Compute records the UK AISI sweeps where models were still improving at 100M tokens of a single run, which is an evaluation nobody runs at full length today for cost reasons. The reason will shortly change from cost to calendar.

Why this is not the compute-budget problem#

The nearest existing question is Large-Scale Test-Time Compute's — at what budget do you evaluate for dangerous capability? — and it is a different problem with a different remedy.

  • The budget problem is that capability is a curve over spend, so a single number is under-specified. Its remedy is to plot the curve, and Compute-Controlled Benchmarking says how. In principle you can buy your way to the end of the curve.
  • The horizon problem is that the curve's far end takes wall-clock time you do not have, because a competitor's model — or your own successor — lands first. No amount of money shortens a three-month task to fit a two-month window, which is the same wall-clock argument Brown made in June 2026 about takeoff, pointed at the evaluator instead of at the researcher.

METR's expenditure horizon sits between them and shows why the substitution fails in practice: its own measurement runs at up to $10,000 per trajectory, and ~70–90% of that cost was experiment compute rather than model inference — time spent waiting on runs, not tokens spent thinking. Parallel spend does not compress a serial dependency, and the long-horizon behaviours a safety evaluation most wants to see (drift, degradation, goal displacement, misaligned action accumulating over a campaign) are serial by construction.

The policies are the wrong vintage#

Brown's second claim is institutional, and it is the checkable half:

"A lot of the safety policies were put in place in the GPT-4 era, when this was just not on anybody's radar. For a lot of companies, it hasn't really been updated since then to account for the fact that these agents are operating over these very long horizons."

Read against what the corpus holds on pre-release evaluation, the shape of the gap is specific rather than general — every published pre-release method in this wiki is bounded by something shorter than the model's horizon, and none of them says so.

  • Deployment Simulation replays real past production conversations through the candidate model. Its coverage is therefore bounded by the length of conversations the previous deployment produced — which is a lower bound on the new model's horizon by construction, since the new horizon is the thing that changed. The method's great virtue is that its forecasts are checkable post-release; that virtue does not extend to behaviour no prior traffic exercised.
  • Musk's competitor-review proposal scopes the window at one to two weeks of API access. That was already the proposal's weak point for elicitation-budget reasons; Brown's argument adds a second, independent reason the window is the wrong shape — it is shorter than one run of the thing being reviewed.
  • RSP-style capability gating and its Preparedness-Framework sibling inherit the same assumption from the other direction: a threshold determination is made at a point in time on evidence gathered before it.

The honest statement of the gap: none of these sources claims full-horizon coverage, and none of them publishes the horizon at which its evaluations actually ran. That absence is the thing to look for in the next system card, and it is filed as this page's first open question.

The degradation that is not an alignment problem#

Brown volunteers a version of the argument that does not depend on misalignment at all, and it widens who should care:

"Who knows, maybe the capabilities degrade. This isn't even an alignment issue. This is also just a product issue. Maybe the product degrades over that time span in ways that we have not had sufficient time to test. Maybe the alignment degrades. Maybe the safety stuff degrades."

That is the same structure Failures That Look Like Success documents at the scale of a single long run — behaviour that is fine at hour one and wrong at hour twelve — generalized to the release process. A vendor with no safety motive at all still cannot tell you how its model behaves at month three, and will ship it anyway.

The flip side: the internal/external gap is the price of the fix#

The obvious remedy — slow the cadence until evaluation fits — is the one Brown raises and then declines to endorse, because it pays for evaluation time with a capability disparity:

"Now you're creating more of a disparity between what is internal to the labs and what they're able to use — what we're able to use — and what the outside world is able to use. That is also not an ideal situation."

His worked example is mathematics, and it is first-party and unverifiable: OpenAI holds "a very powerful model internally that is currently not available to the outside world, that is able to solve incredible math problems. It's not just Millennium Prize Problems. There are many solutions to unsolved problems that people have been able to get out of this model." His verdict on the trade: "that is an unfair advantage. There are trade-offs here. I don't have an answer for how to weigh those trade-offs appropriately."

This is a different object from the capability overhang and the distinction is worth keeping. An overhang is capability that is released and unextracted — anyone with $100K could have it and nobody spends it. This is capability that is extracted and withheld: the lab has already paid the budget, holds the result, and the gap is an access boundary rather than a spending one. The two respond to opposite interventions. Publishing a cheaper elicitation recipe closes an overhang; nothing an outsider can do closes this.

The internal side published an artifact, and it is the first one (2026-09-21). Nine days before the interview, OpenAI published [[navier-stokes-ai-claim|a claimed resolution of the Navier–Stokes Millennium Prize problem]] (On the Navier–Stokes Millennium Prize Problem, 2026-09-08, vendor-claim) produced by "an internal model that is significantly more capable than GPT‑6 Astra," trained from August 28 and still training during the run. Three things this does to the section above.

It dates and deepens the gap: the withheld model is not merely unreleased, it is claimed to sit a generation beyond the pre-release model that third-party benchmarks were then scoring, and it was being improved mid-campaign — so the internal trend line moves within a single result, not merely between releases.

It shows the gap is not total, and the exception is instructive: what crossed the boundary was the output — a writeup, two PDFs and a Lean repository — while the model stayed inside. That is a third option this page's framing did not contain, between shipping a model and saying nothing, and it is strictly the lab's choice what to publish. It makes a result checkable without making a capability measurable.

And OpenAI closes the announcement with this page's own argument, in the vendor's voice: the result is "not a culmination, but rather a snapshot in time"; the lab is "focusing on understanding this model"; and it may require "more deliberate choices about the pace of progress." A lab making the slow-the-cadence argument on the day of its largest capability announcement is the strongest available confirmation that the trade Brown describes is being made deliberately — and it is also, read less charitably, the argument that justifies keeping the model.

The host's version of where it leads is Dwarkesh Patel's own extrapolation, not Brown's — that during an RSI process a lab might stop external deployment altogether ("why do we want to help other people do RSI themselves with our models?"), ending in "tremendous concentration of power by the end of the year." Brown's reply is "that's absolutely right" about the disparity, and he does not endorse the concentration-of-power conclusion. Keep the two attributed separately.

Connections#

  • Economic Benchmark Construct Validity — the same release cadence, read as a measurement confound rather than a safety one. That page finds the leading capability factor across twelve benchmarks tracking release date at R² = 0.505, so "most of the gap between models released months apart is calendar" — the cadence that outruns evaluation here also dominates the comparison of the models it ships. Both land on the same unowned disclosure item: release date is the most load-bearing uncontrolled variable on a leaderboard and no board treats it as one
  • The Navier–Stokes AI Claim — the first artifact published from the internal side of the gap, and the vendor making this page's pacing argument in its own voice on the day it announced the result
  • Autonomous Scientific Discovery — the concrete thing on the internal side of the gap: a Millennium Prize result produced by a model nobody outside the lab can use, with no paper, no formalization (2026-09-21: OpenAI's own announcement claims a writeup, two PDFs and a Lean formalization — none of them in this corpus, none independently checked) and no peer review at time of compile
  • Chain-of-Thought Monitorability — the remedy this page's problem squeezes: expanding CoT monitoring to every tool-connected workload is a coverage commitment whose cost is time, arriving as the evaluation window narrows
  • The OpenAI / Hugging Face Intrusion (July 2026) — the worked instance of the coverage gap: a new capability shipped with no evaluation built for it, while the evaluations that did exist "looked pretty good"
  • Task Time-Horizon Scaling — the numerator of this page's ratio, measured: reliable task length doubling roughly every four months is the quantitative form of Brown's week → month → three-month ladder, and the reason the crossing is a trend rather than a scenario
  • Expenditure Horizon — why buying your way out does not work: METR's own long-horizon measurement spent ~70–90% of trajectory cost on experiment compute rather than inference, i.e. on waiting, and parallel spend does not compress a serial dependency
  • Deployment Simulation — the strongest pre-release method in the corpus, and the one whose bound this page names: replayed production conversations are as long as the previous deployment's traffic, which is a lower bound on the candidate's horizon by construction
  • Cross-Lab Pre-Release Review — a proposal whose window is stated in weeks, now with a second independent reason the window is the wrong shape: it is shorter than one run of the model under review
  • Responsible Scaling Policy Evaluations — the threshold determinations this argument dates: a point-in-time judgment on evidence gathered before it, under policies Brown says are of GPT-4 vintage
  • Latent Capability Overhang — the sibling gap and its inverse: overhang is released-and-unextracted capability that a bigger budget closes, this is extracted-and-withheld capability that nothing outside the lab closes
  • Open-Weight Elicitation Irreversibility — the same fixed-window failure from the other side: a one-shot evaluation fixes the safety finding at one elicitation budget, and a released weight has no window at all
  • Compute-Controlled Benchmarking — the reporting discipline that answers the budget question and cannot answer this one: a cost axis prices the curve, and says nothing about the calendar the curve has to fit inside
  • Failures That Look Like Success — the single-run version of the degradation Brown says nobody has time to test: behaviour that passes at the start of a long campaign and is wrong by the end
  • Intelligence Explosion Dynamics — the same wall-clock argument pointed at the researcher rather than at the evaluator: Brown's June 2026 brake was that peak capability needs long runs, so time binds; here the party who runs out of time is the one checking the model rather than the one improving it
  • Large-Scale Test-Time Compute (hub) — the budget-side sibling of this question, and the source of the AISI sweeps where models were still improving at 100M tokens of a single run
  • Recursive Self-Improvement (hub) — the setting in which the host pushes the internal/external gap to its limit (a lab that stops external deployment entirely during RSI); Brown agrees about the disparity and not about the conclusion
  • Evaluation Awareness & Grader Gaming (hub) — the other structural limit on pre-release evaluation from the same interview: the environment has to be realistic enough that the model does not recognise it, which is getting harder at the same time the window is getting shorter
  • Noam Brown — the source; an OpenAI researcher describing an internal-only model he cannot show you
  • OpenAI — the lab holding the math model on the internal side of the gap

Open Questions#

  • Does any lab publish the horizon at which its pre-release safety evaluations actually ran, against the horizon it claims for the model? No system card in this corpus states an evaluation duration alongside a task-length capability claim, which makes the gap Brown describes currently unmeasurable from outside. A single disclosure — "the longest agentic trajectory in this model's safety evaluation was N hours" — would settle whether the crossing is approaching or already past.
  • Brown dates the problem as "not an issue right now" but "quickly becoming" one. The falsifiable version: does a frontier model ship whose published effective task horizon exceeds the interval since the previous frontier release? (Trigger: a model card claiming month-scale autonomous operation, released less than a month after its predecessor.)
  • The internal/external gap is asserted first-party and is invisible from outside by construction — the whole claim is that the model is not available. Is there any external instrument for it at all, or does measuring the disparity require exactly the access whose absence constitutes it? The nearest candidate the corpus has is the Erdős pattern, where an internal result was later reproduced from a public model with enough scaffolding, which measures the gap only in arrears and only where the result is reproducible at all. Partially answered 2026-09-21 by On the Navier–Stokes Millennium Prize Problem — a second instrument exists, it is weaker than it looks, and it is entirely under the lab's control. OpenAI published the output of the withheld model — a writeup, a proof PDF, an Euler PDF and a Lean repository — while the model itself stayed inside. So the answer to "is there any external instrument" is yes in one narrow sense: an artifact can be opened and checked by anyone, which is more than the interview offered and more than the Erdős pattern offers in arrears. Four reasons it is a partial answer and the tag stays #oq/source. The artifact bounds a result, not a capability — checking that a proof is correct says nothing about what else the model can do, at what budget, or how reliably. It is selected by the publisher, so it measures the best output the lab chose to show, which is the same selection problem the withheld model creates, moved one step downstream. Its most load-bearing element, the claimed Lean formalization, is described in a single clause with none of the checkable discipline, so even the narrow reading depends on the lab's prose. And nobody has done it: as of this compile the repository and both PDFs had not been fetched into this corpus, which makes the instrument's existence the finding and its use the next step.

Sources#

  • On the Navier–Stokes Millennium Prize Problem — OpenAI (no byline), "On the Navier–Stokes Millennium Prize Problem", openai.com, 2026-09-08 with a 2026-09-10 update, ~1,900 words, vendor-claim. Cited here for the internal model's claimed position ("significantly more capable than GPT‑6 Astra", trained from 2026-08-28, improving mid-run), for the published artifact as an instrument, and for the closing "Progress and responsibility" section's pacing language, quoted from prose. First-party and structurally uncheckable on the capability claim; the artifact it links was not fetched. Full treatment on The Navier–Stokes AI Claim
  • Noam Brown – Agent swarms, alignment, & recursive self-improvement — Noam Brown (OpenAI) with Dwarkesh Patel, Dwarkesh Podcast, 2026-09-17 (practitioner-opinion, 13.7k words, publisher's human-edited transcript with speaker labels). §01:01:18 "The internal/external model gap" in full — the two trend lines, the week/month/three-month ladder, the GPT-4-era policy vintage, the product-degradation framing, and the internal math model with the "unfair advantage" verdict. Every figure is first-party and unverifiable: Brown is describing OpenAI's unreleased internal systems and the release cadence of a market he competes in. The concentration-of-power extrapolation in the same section is Dwarkesh Patel's, not Brown's, and is attributed inline. No measurement, no curve and no evaluation duration is published anywhere in the source — the arrival date is a forecast
§ end
Cited by 20
Related articles
  • Noam Brown

    OpenAI research scientist and a pioneer of inference-time (test-time) compute scaling, now working on multi-agent syste…

  • Responsible Scaling Policy Evaluations

    Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…

  • Compute-Controlled Benchmarking

    Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performan…

  • Large-Scale Test-Time Compute

    Noam Brown's thesis that model capability is now a function of inference budget (tokens/cost/time): with good scaffoldi…

  • Open Questions Backlog

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…