H
Howardism
Plate IIModel Capability & TrainingHOWARDISM

Diversity Calibration Under SFT

Skobelev, Fithian & Han (arXiv 2609.16454): output diversity is measured as collision probability, and plain SFT is not inherently biased toward mode collapse or over-dispersion — a bias-variance decomposition lets either sign occur, and a sqrt-KL bound forces the gap to zero as the fit improves; survey and CodeNet fine-tunes converge to human-level diversity from both directions

Article metadata
Publication details
Published:September 30, 2026
Filed:Concept
Domain:Model Capability & Training
Tags:SFTDiversityMode CollapseCalibrationEmpirical
Reading:7 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Diversity Calibration Under SFT

Sources#

Summary#

The received view is that low output diversity ("mode collapse") is an inherent limitation of LLMs, cited from story-writing, idea-generation and silicon-sampling studies. Skobelev, Fithian and Han (Northwestern / UChicago, arXiv 2609.16454, empirical) argue it is instead a finite-sample fitting error measured against a target: with enough supervised fine-tuning (SFT) data drawn from the target distribution, model diversity converges to the target's, and before convergence the error can have either sign. Models can be over-dispersed as easily as under-dispersed, depending on model and dataset.

The measurement#

Diversity is collision probability: the chance two independent responses to the same fixed prompt coincide, C(π) = Σ π(a)² for tokens, or the expected similarity E[k(Y,Y')] under a kernel k ∈ [0,1] when exact matches are too rare (embedding cosine, cluster membership, normalized Zhang-Shasha tree similarity on code ASTs). The collision ratio R = C(q)/C(p) compares model q with target p: R = 1 calibrated, R > 1 mode collapse, R < 1 over-dispersion. R = 2 means the model offers half as many effective choices as the data (inverse Simpson index). Conclusions are metric-specific, since form-sensitive and content-sensitive diversity metrics see different things.

Theory: no built-in direction, but a hard bound#

  • Decomposition (Eq. 6). Across independent fits at a context h, E[R] = 1 + (1/C(p)) · [ Tr Cov(q̂) + ‖b‖² + 2pᵀb ]: a nonnegative variance term, a nonnegative squared-bias term (both push toward collapse), and a sign-indefinite target-bias alignment term 2pᵀb that can outweigh them. The alignment term is positive when the starting distribution overlaps the target's modes more than the target overlaps itself, negative when the start is diffuse. Finite samples alone therefore do not imply mode collapse. If bias is zero the model is calibrated or under-dispersed in expectation.
  • Bound (Theorem 1). |C_k(p) − C_k(q)| ≤ TV(p⊗p, q⊗q) ≤ √KL(p‖q) (product-measure total variation plus Pinsker). A model close to optimal under population cross-entropy cannot have arbitrarily miscalibrated diversity, and the kernel effective diversity obeys D_k(q) ≥ D_k(p) / (1 + D_k(p)√ε). The bound holds for any fitting procedure, not only SFT.

Evidence#

Experiment 1, synthetic languages. 100 small GPTs at each of 100 log-spaced sample sizes (10,000 per language) on two order-16 languages with exactly known targets, so each term in Eq. 6 is measurable. From a diffuse prior the cross term is negative at every N, and the smallest N shows over-dispersion before variance and squared bias take over and produce collapse. From a mode-aligned prior the cross term turns positive at intermediate N but accounts for at most 31% of E[R] − 1. The authors flag the two settings differ in start, model size and language, so this is not a clean ablation of the prior alone. The bound held on every evaluated context.

Experiment 2, surveys. LoRA fine-tunes of gemma-2-2b-it, gemma-3-4b-it, Qwen3.5-2B and Qwen3.5-4B on GSS (376 questions), WVS (53) and ANES (44), scored per question-by-demographic-group pair (18,270 / 8,015 / 8,211 pairs, groups with ≥60 respondents). Zero-shot, three of four models have median R of 1.61 to 2.92, roughly one-third to two-thirds of human effective diversity, worst on the sparse cross-national WVS. Beyond about 3,000 fine-tuning examples all twelve model-survey combinations sit within 10% of R = 1. Qwen3.5-2B, the best-calibrated base model, is driven away from 1 early in fine-tuning (to 2.22 on GSS, under half its starting effective diversity) before recovering, so small fine-tuning sets can worsen calibration.

Experiment 3, CodeNet. Same four models on accepted Python submissions (2,647 candidate problems, 32 programs per problem, tree similarity). Base medians fall on both sides of the target: both Qwen models over-diverse, both Gemma models under-diverse (R range 0.74 to 1.64). SFT moves all four medians toward 1, and the result persists on the subset of problems present at both stages.

Scope and limits#

  • The result is about plain SFT on target-distribution data. It says nothing about why preference optimization contracts diversity; the paper cites Kirk et al. (PPO-RLHF reduces diversity on summarization but not instruction following), Karouzos et al. (the collapsing stage varies by lineage: SFT or DPO), and GX-Chen et al. (KL-regularized RL is designed to mode-collapse, so the KL penalty alone does not preserve the reference distribution). The finite-sample gap is additive to that objective-induced gap C(π*) − C(p).
  • The bound requires the model to be close in KL, so it needs sufficient coverage of the target distribution; the conclusion names limited data as the open challenge. All models are 2B to 4B parameters and the targets are human survey answers and code, not open-ended creative text.
  • Diversity is evaluated relative to a specified target, which the authors recommend reporting alongside performance benchmarks. The practical claim for silicon sampling is that a fine-tune on the target population's answers, not a diversity-preserving objective (GEM, SED-SFT, TOFU loss), is enough here; the paper did not compare against those methods.
  • Collapse found in prior studies of zero-shot or instruction-tuned models is not contradicted, only reinterpreted: it is what a model far from the target looks like, not a ceiling.

Connections#

  • Alignment Fine-Tuning (AFT) — the SFT + RLHF stage this paper isolates the SFT half of; its diversity effect depends on the target data, not on SFT as such
  • Trained Calibration — the other sense of "calibration" in this wiki (confidence tracks correctness, trained by proper scoring rules); this page is distributional calibration of output spread, and both are fixed by fitting the right target
  • Design by Selection — the practitioner-side symptom, a homogeneous default aesthetic, that this page re-reads as distance from a target distribution
  • Controlled Variance: AI's Edge as Reduced Dispersion — the same word pair from the other side: there reduced dispersion in an AI's execution is the benefit; here dispersion matching the human target is the goal
  • AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins — the individual-level version of the silicon-sampling question, with no fine-tuning. In-context digital twins of 139 consumers miss in one direction: they are more deliberative than the person they copy (+0.25 to +0.87 on a 0–4 System-1/System-2 scale, in two model families), and feeding them a richer interview transcript does not improve their predictions. This page shows the population spread can be fit; that one shows per-person reasoning style is not fixed by more context

Open Questions#

  • Does the √KL convergence survive RLHF/DPO stages, where the optimum itself is concentrated (GX-Chen et al.)? Does post-SFT preference tuning re-open a collision gap that a larger SFT set cannot close?
  • At what fine-tuning size does calibration arrive for open-ended text (stories, ideas) with a semantic-similarity kernel, given the experiments cover only categorical survey answers and code ASTs at 2B to 4B parameters?

Sources#

  • Fine-Tuning Fixes Mode Collapse and Over-Dispersion in LLMs — Skobelev, Fithian & Han, arXiv 2609.16454 (2026-09-15), empirical, 33 pp: collision-ratio framework, bias-variance decomposition, Theorem 1, synthetic-language, GSS/WVS/ANES and CodeNet experiments. Quoted from prose and captions; Table 2 (survey medians) not cited by row.
§ end
Cited by 7
Related articles