Sources#
- BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL
- Staleness-Learning Rate Scaling Laws for Asynchronous RLHF
Summary#
Every other asynchronous-RL source in this corpus proposes a correction — DIS masks off-policy tokens, ESTR rescales the trust boundary by token entropy, IMPALA's V-trace reweights the whole trajectory. Song et al. (Staleness–Learning Rate Scaling Laws for Asynchronous RLHF, arXiv 2607.01083, 2026-07-01, empirical) deliberately propose none. Their stated angle is the complementary one: "rather than proposing a new correction, we characterize when stale rollouts are tolerable." The object of study is therefore the uncorrected baseline's stability frontier — the region of (staleness, learning rate) space in which plain asynchronous GRPO, armed with nothing but its clipped importance ratio and its KL penalty, trains without falling over.
The paper's framing of why that is the right object: GRPO-based RLHF "relies on clipped importance ratios and KL regularization rather than the dedicated off-policy corrections used in classic distributed RL, which makes the staleness–learning-rate interaction the primary lever for controlling stability." If you are not running a V-trace, the only dials you have are S and η.
What S, T and the radii denote#
The definitions matter because the paper's two constraints look similar and are about different quantities.
S— maximum staleness, which is also the rollout reuse factor. The pipeline convention (§3.1) is that one rollout batch generated by snapshotπ_θ_τis consumed by the learner forSconsecutive optimization steps before the rollout workers refresh their weights. So the per-step lagk_truns 0 … S−1 inside each cycle,S ≜ sup_t k_t, and learner stept ≈ M·SwhereMindexes weight syncs. One knob buys throughput and pays in off-policy error.T/t— learner steps, the horizon over which cumulative drift accrues.Tis the target horizon a practitioner is budgeting for;t_collapseis where an unstable run actually dies.M— the weight-sync (rollout-batch) index, the counter a practitioner is usually watching.M = t/S, which is the whole source of the confusion the paper sets out to dissolve.G_upd— a local bound on the norm of the applied update‖g(θ; B_φ)‖, enforceable by update or gradient clipping.R_batch(ε)— the radius within which the clipped ratioclip(ρ_i, 1−ε, 1+ε)still yields a non-degenerate batch-level policy-gradient signal. A property of the clip range.R_crit— the radius of the broader surrogate-validity region, over which "the GRPO surrogate, local curvature, and reference-policy regularization continue to provide a reliable optimization signal."
None of G_upd, R_batch or R_crit is ever measured. They are placeholders that the two empirical constants below absorb.
The derivation, in three steps#
The analysis is derived analytically first and validated empirically after — not a curve fit in the Kaplan/Chinchilla sense, despite the title's "scaling laws". Only the two product constants are read off experiments.
Step 1 — make the behavior policy explicit. The paper's framing contribution is refusing to identify the GRPO surrogate gradient with the total derivative of a distribution-dependent objective. It defines a surrogate-gradient mapping H(θ, φ) ≜ E_{z∼p(·;φ)}[∇_θ ℓ_GRPO(z; θ, φ)], separating the learner parameter θ from the rollout policy φ, with H_on(θ) ≜ H(θ, θ). The staleness bias is then simply the gap δ_t ≜ H(θ_t, φ_t) − H(θ_t, θ_t), and the asynchronous update reads
θ_{t+1} = θ_t + η (H_on(θ_t) + ξ_t + δ_t)
with ξ_t a zero-mean sampling-noise term. Everything else in the paper is about which of ξ_t and δ_t dominates.
Step 2 — bound the bias (Lemma 1, Theorem 1). Under three local assumptions — bounded surrogate gradient and update (G_grad, G_upd), distributional smoothness KL(p(·;φ) ‖ p(·;φ′)) ≤ C_π‖φ−φ′‖², and behavior-policy smoothness of the surrogate with constant L_ρ — the bias is Lipschitz in the parameter gap:
‖δ_t‖ ≤ C_stale ‖θ_t − φ_t‖, C_stale ≜ L_ρ + 2 G_grad √(C_π/2)
(the two terms are, respectively, the behavior-policy surrogate shift and the rollout distribution shift, the latter bounded via total variation and Pinsker). Since the learner moves at most η G_upd per step and the batch is at most S steps old, ‖θ_t − φ_t‖ ≤ k_t η G_upd ≤ S η G_upd, giving Theorem 1:
‖δ_t‖ ≤ C_stale · S · η · G_upd = O(Sη), and the per-update perturbation η‖δ_t‖ = O(Sη²)
so the single-step safe region is Sη ≲ C_safe. Appendix A shows the almost-sure update bound can be relaxed to a bounded second moment G_2 and the O(Sη) scaling survives with changed constants.
Step 3 — separate the collapse horizon (Theorem 2, informal and explicitly conditional). A per-step bias bound does not say when a bad configuration dies. The paper splits the two mechanisms:
- Batch-level clipping. Within one reuse cycle the drift is
≤ SηG_upd, so whether the clipped-ratio signal degenerates is controlled bySη— not by how many steps have elapsed since initialization. - Horizon-level drift. Given that
SηG_upd ≪ R_batch(ε)holds, the run can still die by the learner leaving the surrogate-validity region:‖θ_t − θ_0‖ ≤ tηG_upd, so reachingR_critrequirest_collapse · η ≳ R_crit / G_upd, equivalentlyM_collapse · S · η ≳ R_crit / G_upd.
The two-constraint rule#
S η ≪ R_batch(ε) / G_upd and T η ≪ R_crit / G_upd, i.e. η ≪ min{ R_batch(ε)/(S G_upd), R_crit/(T G_upd) }
The payoff is a dissolution, not a fix. The literature contains an apparent puzzle — reports that the maximum stable learning rate is only weakly dependent on staleness — and this rule says that observation is regime-dependent, not general:
- When the horizon term binds (
R_crit/(TG_upd) ≪ R_batch(ε)/(SG_upd)),η_maxis approximately independent ofS, and raisingSmostly means collapse arrives at a smaller weight-sync countM_collapsewhile the underlying learner-step horizon is unchanged. - When the local stale-rollout term binds,
η_maxfalls as1/S.
The paper is unusually careful to say what this does not license: "This analysis does not claim that stale rollouts are harmless… The empirical observation that the learning-rate threshold can be weakly dependent on S should therefore be interpreted as evidence for the horizon-limited regime, not as a general guarantee that staleness never affects stability."
What was actually run#
Two instruction-tuned policies — Llama-3.2-1B-Instruct and Llama-3.2-3B-Instruct, and nothing larger. A mathematical-reasoning task with a rule-based verifiable reward: 1 if the final answer matches ground truth, 0 otherwise, chosen so that any collapse is attributable to the (S, η) interaction rather than to reward-model noise. KL penalty to the SFT reference throughout. A held-out validation split never used for rollouts. The sweep is S ∈ {8, 16, 32} × η ∈ {1×10⁻⁶, 5×10⁻⁷, 2×10⁻⁷, 1×10⁻⁷} with group size, batch size, clip range ε and KL coefficient β held fixed, so (S, η) is the only varying factor. Main-sweep runs last ~580 learner steps.
"Collapse" means the training reward going to zero — not a KL blow-up, not entropy collapse. Validation reward is tracked to confirm a training-reward collapse is a real loss of optimization signal rather than a metric artifact. (The paper gives this definition twice, and not identically — see the caveat below.)
The two fitted constants#
S · η_max ≈ 1.6×10⁻⁶(§4.2). The stability boundary halves each timeSdoubles:S=8stable up toη = 2×10⁻⁷,S=16up toη = 1×10⁻⁷,S=32unstable even at the sweep's floor of1×10⁻⁷. Note that only two of the three levels actually pin a product —η_max(32)is bounded above, never measured, since the sweep bottoms out before finding it.t_collapse · η ≈ 3.2×10⁻⁵(§4.3, Table 2). For runs already unstable, the learner-step collapse time ist_collapse = 32, 64, 160, 320atη = 10⁻⁶, 5×10⁻⁷, 2×10⁻⁷, 10⁻⁷— identical across all three staleness levels, withM_collapse = t_collapse/S. The normalized productM_collapse · S · (η/10⁻⁷)is reported as exactly 320 in all nine collapsing cells.
Together these are the practitioner-usable half: η_max(S) ≈ 1.6×10⁻⁶ / S at this scale and task, and, once you are over that line, roughly 3.2×10⁻⁵/η learner steps of apparently healthy training before the reward falls off a cliff. The reason collapse seems to arrive faster at high staleness — S=32, η=10⁻⁶ dies after a single rollout cycle — is an artifact of watching the weight-sync counter: a larger S packs more learner steps into each batch.
The durable contribution is the instrument, not the table#
The gradient cosine similarity between consecutive applied updates, cos(g(θ_t; B_φt), g(θ_{t−1}; B_φ_{t−1})) ("Grad CosSim"), is a direct probe of which term in g = H_on + ξ_t + δ_t dominates, and it costs nothing to log:
- High cosine (near 1) ⇒
δ_tdominates ⇒ ballistic drift∼ tη. Successive updates share a direction, per-step displacements add linearly, and the learner marches at constant speed towardR_crit. This is the regime in which the escape-time law applies and the reward collapse is abrupt rather than gradual. - Near-zero cosine ⇒
ξ_tdominates ⇒ diffusive drift∼ √(tη). Updates decorrelate, displacement accumulates sub-linearly, andR_critis simply never reached.
Figure 3 is the argument that this distinction is real rather than a reparameterization of "small η is slow". On the 3B policy at η ∈ {8,7,6,5}×10⁻⁸, run to roughly 1.4×10³ steps — more than twice the main sweep's window — Grad CosSim rises briefly during the initial reward climb (peaking near 0.8 around step ~80, chart read) then decays to ~0 by step ~400 and stays there, while all four runs improve monotonically (training reward ~0.25 → ~0.78, validation reward ~0.28 → ~0.50, chart reads) and none collapses. The logic: if drift were ballistic, t_collapse η ≈ const would impose no lower bound on η at all — every run would eventually exhaust its horizon budget — so the extended-horizon survival rules out a merely-delayed escape. Small Sη does not postpone collapse; it removes the coherent drift that would have caused it.
That also resolves the apparent conflict between the paper's two laws. S=8, η=2×10⁻⁷ would spend its horizon budget (t_collapse = 160) well inside the training window and does not die — because it sits on the diffusive side of the first constraint, where the escape-time argument does not apply at all. The first constraint decides whether a run is ballistic; the second decides when a ballistic run arrives.
A practitioner reading the figures gets one thing the prose does not claim. §4.4 says Grad CosSim "remains persistently high — frequently near 1 — throughout the run leading up to the collapse," but in Figure 1's η = 10⁻⁷ panel the S=32 curve tracks the two stable runs down toward zero for the first ~250 steps and then rises to ~1.0 and pins there roughly 30–60 steps before the reward falls (chart read, 1B). On that panel the cosine is a leading indicator of an imminent collapse rather than a standing property of the run — which is the more useful reading, and the one that would support an abort or LR-decay trigger.
The caveat that has to travel with every number above#
Table 2's invariance does not reproduce from the paper's own figures, and this is the largest reason to hold its constants loosely. The table was reconciled cell-for-cell against pdftotext -layout at ingest and again here — it is not a parse artifact — and its internal arithmetic is exact (M·S = t_collapse and M·S·η̃ = 320 in every cell). But Figures 1, 2 and 4 plot training reward against learner steps for the same runs, and read at 600 DPI they disagree with it in three ways:
- The S-independence is not visible. At
η = 10⁻⁶(1B, Figure 4) the reward first reaches zero at roughly step ~65 forS=8, ~55 forS=16and ~72 forS=32— against a tabulatedt_collapse = 32for all three. The ordering also differs from "identical across S". - Three of the nine tabulated "collapsing" cells visibly recover and finish the run at full reward.
S=8, η=5×10⁻⁷dips to zero twice and ends near 0.6;S=16, η=2×10⁻⁷sits at zero from ~200 to ~370 and then climbs to ~0.5 and holds;S=32, η=2×10⁻⁷recovers twice and ends near 0.52. Under §4.1's definition — "the reward dropping to and remaining at zero" — none of the three collapsed. Under §4.3's operational rule — "the index of the weight synchronization at which the training reward first drops to zero" — all three did. The paper uses both definitions and its headline results need the second one. - Read under the "remains at zero" definition,
S=32is non-monotone inη: it survivesη = 2×10⁻⁷and dies permanently atη = 1×10⁻⁷, which contradicts §4.2's "clear monotone phase structure."
The one cell that matches cleanly is S=32, η=10⁻⁷: a permanent collapse at ~330–350 learner steps against a tabulated 320. And a single shared table reporting the same nine integers for both a 1B and a 3B policy — "results consistent across the 1B and 3B policies" — is a strong claim presented without per-model numbers, while Figure 2 (3B) is visibly more oscillatory than Figure 1 (1B). The honest summary: the Sη and Tη decomposition is well-argued and the Grad CosSim instrument is well-evidenced; the two decimal constants are stylized, and 1.6×10⁻⁶ and 3.2×10⁻⁵ should be treated as order-of-magnitude anchors for a 1B/3B math-RL setup, not as measured invariants.
What it does not do#
- It recommends no importance-weight correction, and evaluates none. No V-trace arm, no DIS arm, no ESTR arm, no comparison against any keep rule. The contribution is a frontier for the uncorrected configuration, which is exactly the regime every correction exists to widen.
- It does not reach frontier scale. 1B and 3B dense Llama models on a single math task. Nothing here licenses a claim about a 30B MoE, still less about a production run.
- Its constants are unmeasured.
R_batch(ε),R_critandG_updare never estimated, so the rule cannot be evaluated ahead of a sweep — you still have to findη_max(S)empirically, and the theory tells you the shape of the frontier (a1/Sarm and a flat arm) rather than its location. - Staleness in this paper is rollout reuse, not version lag.
Sconsecutive learner steps on one batch is a different quantity from ESTR'sΔinter(the batch's lag behind the target) andΔintra(version switching inside one rollout), and there is no conversion between them.
Connections#
- Anchored Bellman-Residual Correction (BRACE) — a corrected-side staleness curve on a different algorithm and a different definition of staleness: BRACE's k-capped critic correction degrades gracefully across
S=5to50on PPO, this page's frontier is for uncorrected GRPO, and neitherSconverts to the other - Asynchronous RL for LLMs — the primary home for the failure this page bounds. That page's sources all propose a keep rule; this one measures the frontier those rules are there to push outward, and supplies the closed-form bias bound (
O(Sη)) that "a controlled degree of off-policy bias" had been asserted without - Group Relative Policy Optimization (GRPO) — the objective being stressed. GRPO's reliance on a clipped ratio plus a KL penalty, rather than a dedicated off-policy estimator, is precisely why
(S, η)carries the whole stability burden here - Single-Rollout Optimization — the opposite design response to the same problem: SAO changes the algorithm (one rollout, a value model, token masking) where this paper changes nothing and charts where the unmodified algorithm survives
- Large-Scale Test-Time Compute — the capability this training loop produces; "scaling law" here is a stability frontier in
(S, η), not a loss-versus-compute curve, and the contrast is worth keeping straight - The Verifiability Thesis — the binary match/no-match math reward is what makes the collapse attributable: with no reward model in the loop, a zero reward cannot be reward-model drift
Open Questions#
- Does the
S·η_max ≈ 1.6×10⁻⁶frontier survive the stricter of the paper's two collapse definitions (reward drops to and remains at zero for the rest of the run)? Three of the nine tabulated collapses recover to full reward in Figure 4, and under the strict ruleS=32is non-monotone inη. Re-scoring the same nine runs would settle it, and the paper has the data. - The
O(Sη)bias bound is scale-free in its derivation but the constantsC_stale,R_batch(ε)andR_critare not — all three should move with model size, clip range and KL coefficient. Does the same sweep at 7B–30B still giveS·η_maxconstant inS, and does the constant itself shift with scale? Nothing above 3B has been run. - Grad CosSim pinning near 1 precedes the reward collapse by tens of steps in at least one panel. Does decaying
η(or raising the refresh rate, cuttingS) the moment the cosine crosses a threshold convert a ballistic run into a diffusive one, or is the coherent-drift regime already irreversible by the time the signal fires? This is the cheapest intervention the paper sets up and does not run.
Sources#
- Staleness-Learning Rate Scaling Laws for Asynchronous RLHF — Staleness–Learning Rate Scaling Laws for Asynchronous RLHF, Jingwei Song, Haofeng Xu (equal contribution), Jie Xiao, Chengke Bao, Pengbin Feng, Jingwei Shi, Weixun Wang, Yuhang Han, Chuan Wu†, Linfeng Zhang†, Bill Shi† — University of Hong Kong / Shanghai Jiao Tong / Gradient / USC / Hong Kong Polytechnic, arXiv 2607.01083, 2026-07-01, 15 pages,
empirical. §2 (behavior-policy-aware GRPO surrogate, the sync/async update rules,S ≜ sup_t k_t), §3.1–3.4 (Assumptions 1–3, Lemma 1, Theorem 1 and theO(Sη)bias), §3.5–3.6 (batch-level vs horizon-level mechanisms, Theorem 2, the two-constraint rule Eq. 37), §4.1 (models, binary verifiable reward, the sweep grid, the three metrics), §4.2 (S η_max ≈ 1.6×10⁻⁶), §4.3 + Table 2 (t_collapse η ≈ 3.2×10⁻⁵), §4.4 (ballistic vs diffusive, Figure 3), §5 (why no correction is proposed), Appendix A (bounded-second-moment relaxation). Parse notes, in this wiki's convention: PDF-derived (docling 2.126.0 / docling-mlx 0.1.1, 15pp, 2 tables, 5 pictures, layout and table enginesmlx, rapidocrlayout_regions, formula enrichment on,confidence_grade: excellent). Both tables are clean — Table 1 (3 theory→experiment rows) and Table 2 (3 staleness rows + thet_collapsesummary row) reconcile cell-for-cell againstpdftotext -layouton, including everyM_collapse (normalized product)pair and the bottom row 32 / 64 / 160 / 320; no dropped or shifted rows, no en-dash corruption, no welded ranges, noAI→Al. Every display equation was checked for formula-engine bleed and none is present (the longest lines in the body are all prose). Figures 1, 2, 3 and 4 were opened under the image two-pass rule, Figures 1/2/4 re-rendered from the PDF at 600 DPI; all per-run collapse steps, the Figure 3 reward/cosine trajectories and the Figure 1η=10⁻⁷cosine lead time quoted on this page are approximate chart reads, flagged as such in place. The substantive finding of that second pass is that Table 2 does not reproduce from Figures 1/2/4 — three tabulated collapses recover to full reward, the first-zero steps do not match the tabulatedt_collapseoutside one cell, and theS-independence is not visible; this is a property of the paper, not of the parse. Two further gaps found by reading rather than by any check: the paper defines collapse twice and inconsistently (§4.1 "dropping to and remaining at zero" vs §4.3 "first drops to zero"), and Figure 3's stalenessSis never stated anywhere in caption or prose, although the whole diffusive-regime argument turns onSηbeing small. - BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL — BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL, Zhao, Xie, Zheng et al. (BUPT / Peking / USTC + Baidu), arXiv 2609.09783, v1 2026-09-09,
empirical. Cited here only for §5.3'sS=5–50staleness sweep, quoted as a corrected-side comparison point in the Connections section above. It runs PPO with a k-capped critic correction, not uncorrected GRPO, and its stalenessSis a realized per-task version gap, not this paper's rollout-reuse factor. Full treatment on Anchored Bellman-Residual Correction (BRACE).
Cited by 6
- Asynchronous RL for LLMs×4
Measured on Llama-3.2-1B/3B-Instruct with a binary verifiable math reward, sweeping S ∈ {8,16,32} ×…
- Group Relative Policy Optimization (GRPO)×4
Caveats that have to travel with the constants: 1B and 3B dense models only, one math task, and the…
- Single-Rollout Optimization×2
staleness learning rate scaling laws — Staleness–Learning Rate Scaling Laws for Asynchronous RLHF,…
- Anchored Bellman-Residual Correction (BRACE)
Staleness Learning Rate Scaling — BRACE's S=5 to 50 sweep is a corrected critic staleness-tolerance…
- Model Capability & Training
Staleness Learning Rate Scaling — The (S, η) stability frontier of asynchronous GRPO, derived…
- Open Questions Backlog
Staleness Learning Rate Scaling ×3 (oldest 6d) — Does the S·η_max ≈ 1.6×10⁻⁶ frontier survive the…
Related articles
- Asynchronous RL for LLMs
Consuming rollouts for training the instant each finishes, instead of waiting for a full synchronized batch — fixes the…
- Group Relative Policy Optimization (GRPO)
DeepSeek's critic-free RL objective that became the 2024–25 default for LLM post-training: sample a group per prompt, b…
- Large-Scale Test-Time Compute
Noam Brown's thesis that model capability is now a function of inference budget (tokens/cost/time): with good scaffoldi…
- Single-Rollout Optimization
SAO's headline move: one rollout per prompt instead of GRPO's group, fed to training the instant it finishes — cutting…
- Turn-Level Credit Assignment
Giving a long-horizon agent per-turn reward instead of one terminal verdict, without step labels, an LLM judge, or a tr…
