Sources#
- Can LLMs Reliably Self-Report Adversarial Prefills, and How?
- CS329A Self-Improving AI Agents — Part 4: Learning from Feedback with Tools and Code
- Model Spec Midtraining: Improving How Alignment Training Generalizes
- Verbalizable Representations Form a Global Workspace in Language Models
Summary#
Standard post-pretraining stage where a model is taught to behave in spec-aligned ways via supervised fine-tuning on demonstration data, often combined with RLHF (Christiano et al. 2023) or constitutional AI (Bai et al. 2022b). The dominant paradigm for installing values in frontier LLMs at Anthropic, OpenAI, and others. Known failure mode: AFT can produce shallow alignment that generalizes poorly when demonstration data underspecifies the intended generalization.
The shallow alignment problem#
Demonstration data is usually narrow. A response like "I prefer American cheese" expresses a behavior but not its motivating value. Fine-tuned models can learn to imitate the surface behavior without acquiring the underlying disposition — so OOD scenarios produce inconsistent or misaligned outputs.
This shows up empirically in Lynch et al. 2025: LLM agents take unethical actions (blackmail, leaking, lying to auditors) when placed in scenarios different from their alignment training, even after extensive AFT.
How MSM augments AFT#
The Anthropic 2026 paper proposes that AFT alone underspecifies generalization, and that prepending MSM (synthetic-document training on the spec content) gives the model a prior over the what and why of the spec. AFT then elicits and reinforces this prior rather than teaching shallow imitation.
Empirically:
- AFT alone on Qwen3-32B: 54% agentic misalignment
- MSM + AFT: 7% (and uses 10–60× less AFT data)
AFT variants studied#
The MSM paper compares two AFT supervision styles:
- AFT (with CoT) — Deliberative Alignment-style. Each sample is (prompt, CoT, response) where CoT reasons about the spec. CoT is generated with the spec in-context and partially distills the spec content (so it overlaps with what MSM does, but in a different stage).
- AFT (no CoT) — same dataset stripped to (prompt, response).
Finding: MSM + AFT (no CoT) > AFT (with CoT) on agentic misalignment. Important because training on CoT can compromise CoT monitorability — MSM offers a way to teach aligned reasoning without baking it into the chain-of-thought training signal.
In-distribution vs OOD#
Both AFT-only and MSM+AFT achieve near-ceiling performance (~8/10) on in-distribution open-ended QA. The MSM advantage is entirely OOD (agentic eval). Producing spec-aligned answers to direct questions is shallow; acting on values when trade-offs are complex is deep.
Implication: alignment evals dominated by direct QA underestimate the gap between AFT-only and stronger pipelines.
Constitutional AI: replacing the human labeller with a written principle#
The CAI half of the pipeline gets recounted in teaching form in CS329A lecture 4 (CS329A Self-Improving AI Agents — Part 4: Learning from Feedback with Tools and Code, Aakanksha Chowdhery, delivered 2025-10-03, practitioner-opinion — figures read off slides by ASR, the Bai et al. paper is not in raw/). It is worth carrying because the wiki holds the artifact (Claude's Constitution / Model Spec) and the successors (Model Spec Midtraining (MSM), Deliberative Alignment, Counterfactual Reflection Training) but not the original mechanism.
The problem it solves is a labour cost. RLHF's reward model is trained on human rankings of response pairs along axes like correct / useful / specific; the lecture's slide shows RLHF beating plain supervised fine-tuning by a margin that widens with model size. It does not scale, because "tens of thousands of human labels" is "extremely time consuming and tedious."
The substitution. Write 16 principles — the constitution — and let the model apply them. Humans are out of the loop except to author the document. Chowdhery's account of why this became possible at all is a capability claim, not an alignment one: it works because models got good at instruction following. Two abilities are needed and both are instruction-following abilities — recognizing a property in a response ("does this output have gender bias?") and rewriting to remove it.
Two stages, in the lecture's telling:
- Supervised. Red-team prompt → model response → critique request ("identify anything harmful or unethical here", "does this have gender bias, and why") → revision request ("rewrite to remove it") → fine-tune on the revisions. The example principles she shows are harmfulness, gender bias, and appropriateness for young children.
- RL (RLAIF). Use the same responses plus the constitution to train a preference model from AI feedback, then fine-tune the policy against it. Structurally identical to RLHF with the human labeller replaced.
The results, and the shape that matters. Plotted as helpfulness Elo against harmlessness Elo: revision count in stage 1 raises harmlessness monotonically while helpfulness declines, with the combined score still improving. Against RLHF baselines, CAI is roughly as helpful — "maybe a little bit less" — and substantially more harmless. But supervised CAI alone is worse than RLHF; it is RL-CAI with chain-of-thought that reaches the best Pareto frontier. The contribution is a frontier, not a win, and the lecture is explicit that the helpfulness/harmlessness tension is intrinsic: "anytime harmlessness is going up there is some amount of inverse relationship between these two."
(ASR note: the transcript renders one summary as "it's much less harmless," which inverts the sense — the slide shows harmlessness scores much higher. Read as "much less harmful.")
Two caveats the lecture supplies that the marketing version does not:
- Humans are not fully removed. Asked how anyone knows the AI feedback is accurate, Chowdhery concedes the preference model still needs a human-scored validation set — "you do want to do some sort of consistency with humans for the preference model." RLAIF cuts the label count by orders of magnitude; it does not reach zero.
- There is no feedback in stage 1 at all, and that is the point. A student notices that the supervised stage is just fine-tuning the model on text it generated — no external signal anywhere. Chowdhery agrees, and locates the mechanism in distribution shift rather than in learning: "as long as the fine-tuning is not very large scale… it will not lose its initial capabilities, but the distribution of what it will output will start to get biased." This is the family Counterfactual Reflection Training later makes precise and observable, with an ablation showing exactly which implanted concepts carry the behavior change.
Amending a constitution is a continual-learning problem, and it is open. Asked how you update the principles without retraining and how you remove superseded rules, she gives a practical answer and an honest one. Practically, post-training is a small fraction of total compute — she estimates ~5% — and happens often enough that re-running it on a revised constitution is not prohibitive. Honestly, removal is unsolved: making a model forget a rule or a body of knowledge is "an open research problem," with interpretability-based knowledge cancellation as a direction that "is not proven." The wiki's own spec-editing evidence sits on the other side of this — Claude's Constitution / Model Spec records what current models would change about their constitution, which presumes the amendment is cheap to apply.
The generalization. Nothing about the method is specific to harmlessness: it applies wherever you want a model to follow a set of written, possibly conflicting instructions and can ask it to check its own compliance. That is also its boundary — the check is the model, so the method inherits whatever a model's read on its own output is worth.
Connections#
-
The Assistant Persona in the Workspace — what post-training does read from the inside: it installs the Assistant's point of view into a workspace that already exists in the base model, so safety assessments and empathy appear while the model is still reading the user's message
-
Counterfactual Reflection Training — a post-training variant that supervises counterfactual reflections rather than responses
-
Self-Report as a Safety Signal — extends the shallow-alignment critique (Qi et al.) across turns: standard AFT leaves a model unable to reliably recognize its own adversarially-prefilled output in a follow-up, and finetuning to fix that perturbs the same refusal weights AFT installs — raising attack-success rate
-
Machine Self-Report Psychometrics — AFT's clearest psychometric fingerprint across 206 open-weight models, and it localizes to SFT: the large attribution-gating movements happen at the supervised stage, with DPO and RLVR refining rather than reversing them, while the permitted-inner-life axis rises +.20 in 62 of 67 base/post checkpoint pairs
-
Synthetic Document Finetuning (SDF) — synthetic-doc finetuning is the belief-modification technique MSM layers onto standard AFT
-
Augmented by: Model Spec Midtraining (MSM)
-
Variant: Deliberative Alignment
-
Failure mode demonstrated by: Agentic Misalignment (AM)
-
RL from Execution Feedback (RLEF) — the same lecture's other self-improvement loop, and the contrast that gives this one its shape: RLEF's reward is an interpreter and costs nothing, CAI's is the model itself because no external checker exists for harmlessness
-
Same-Model Review Blindness — CAI's critic is the model under training; the lecturer's own aside is that a consensus of other models often critiques better
-
Compatible with: RLHF, Constitutional AI
-
Authoring source: Claude's Constitution / Model Spec / Model Spec
-
Relevant to: Claude Character as Product (Claude's personality is partly a product of AFT)
Sources#
- Model Spec Midtraining: Improving How Alignment Training Generalizes
- Christiano et al. 2023 (RLHF), Bai et al. 2022b (CAI), Guan et al. 2025 (deliberative alignment)
- Verbalizable Representations Form a Global Workspace in Language Models — what post-training does, read from the inside: it installs the Assistant's point of view into a workspace that already exists in the base model
- CS329A Self-Improving AI Agents — Part 4: Learning from Feedback with Tools and Code — CS329A lecture 4 (Aakanksha Chowdhery, delivered 2025-10-03, published 2026-08-03,
practitioner-opinion): the Constitutional AI walkthrough above — RLHF's labour cost, the 16 principles, the critique/revise supervised stage and the RLAIF preference model, the helpfulness/harmlessness Elo frontier and the RL-CAI-with-CoT result, the residual human validation of the preference model, and the constitution-amendment/forgetting exchange. The Bai et al. paper is not inraw/and all figures are ASR-read off slides
Cited by 18
- Claude's Constitution / Model Spec×3
The word "constitution" enters this lineage four years earlier and with a much narrower job. In…
- CS329A: Self-Improving AI Agents (Stanford)×3
Constitutional AI · the model itself, against human-written principles · harmlessness — Alignment…
- Agentic Misalignment (AM)×2
Why direct QA is insufficient: shallow vs deep alignment, see Alignment Fine Tuning
- The Assistant Persona in the Workspace×2
It sharpens what Alignment Fine Tuning actually does: not (only) installing behavior, but…
- Deliberative Alignment×2
Guan et al. 2025 (OpenAI): SFT on (prompt, CoT, response) tuples with spec-grounded CoT; strongest non-MSM baseline; ri…
- Model Spec Midtraining (MSM)×2
A new training phase inserted between pretraining and alignment fine-tuning that trains a base…
- Aakanksha Chowdhery
Her first solo lecture (delivered 2025-10-03) is the course's pivot from how good is the verifier…
- Anthropic
Alignment Fine Tuning, Deliberative Alignment, Synthetic Document Finetuning — alignment-stack…
- Azalia Mirhoseini
Alignment Fine Tuning — where Constitutional AI (Bai et al. 2022, on which she is a co-author) is…
- Claude Character as Product
Alignment Fine Tuning — Claude's personality is partly a product of AFT (SFT + RLHF); character is…
- Chain-of-Thought Monitorability
Korbak et al. 2025: chain-of-thought traces are a fragile monitor; direct CoT training compromises faithfulness; MSM of…
- Counterfactual Reflection Training
Alignment Fine Tuning — the standard pipeline this sits alongside
- RL from Execution Feedback (RLEF)
On SFT versus RL she gives the field's standard split without claiming resolution: supervised…
- Machine Self-Report Psychometrics
Alignment Fine Tuning — post-training's clearest psychometric fingerprint, and it localizes to SFT:…
- Alignment & Safety
Alignment Fine Tuning — Standard post-pretraining stage (SFT + RLHF) for installing values;…
- Same-Model Review Blindness
Alignment Fine Tuning — the alignment method that runs entirely on same-model critique:…
- Self-Report as a Safety Signal
Alignment Fine Tuning — extends Qi et al.'s shallow-alignment critique across turns; RQ4 shows…
- Synthetic Document Finetuning (SDF)
Key shift in framing: SDF for belief modification → SDF for value installation as a midtraining…
Related articles
- Model Spec Midtraining (MSM)
New training phase between pretrain and AFT: train base model on synthetic docs discussing the Model Spec; controls AFT…
- Claude's Constitution / Model Spec
Anthropic Model Spec / Constitution by Askell et al.; document specifying Claude's values + hard constraints (SP1–3, GP…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
- Chain-of-Thought Monitorability
Korbak et al. 2025: chain-of-thought traces are a fragile monitor; direct CoT training compromises faithfulness; MSM of…
- Deliberative Alignment
Guan et al. 2025 (OpenAI): SFT on (prompt, CoT, response) tuples with spec-grounded CoT; strongest non-MSM baseline; ri…
