H
Howardism
Plate IIInterpretability中文HOWARDISM

The Global Workspace in Language Models (J-space)

PublishedJuly 11, 2026FiledConceptDomainInterpretabilityTagsInterpretabilityAlignmentCognitionRepresentationsReading35 minSourceAI-synthesised

Anthropic's July 2026 finding that LLMs maintain a small privileged set of verbalizable representations — the J-space — that satisfies the functional criteria of a cognitive global workspace: verbal report, directed modulation, internal reasoning, flexible generalization, and selectivity; it carries <10% of activation variance and ~25 concepts at a time, yet the causal effects concentrate almost entirely in it

Illustration for The Global Workspace in Language Models (J-space)

Sources#

Summary#

Using the Jacobian Lens (J-lens), Anthropic's interpretability team searched for representations that are verbalizable — poised to be spoken about if the model were asked — and discovered, "rather surprisingly," that this set does far more than support speech. It behaves like a global workspace: a small, privileged, broadcast subset of the model's representations, sitting atop a much larger volume of automatic processing that the model cannot report on or flexibly reason with.

The J-space is the set of points expressible as a sparse non-negative combination of J-lens vectors (typically $k \le 25$). It is small: across workspace layers it never carries more than 10% of activation variance, and a concept's J-space component holds a median 6–7% of that concept's representational variance. And yet that sliver is where the causation lives.

The five functional properties#

Global workspace theory (Baars, Dehaene) says conscious access is reportable, top-down controllable, the medium of deliberate reasoning, flexibly generalizable, and selective. The paper defines an LLM analog of each and tests it.

1. Verbal report. Ask Sonnet 4.5 to think of a sport and read the lens at the colon before it answers: Soccer sits at the top and the model says "Soccer." Swap the Soccer coordinate for Rugby — a word not in its top-10 outputs — and it reports "Rugby." Systematically, swaps drive the target into the top-5 on 88% of trials. Injecting a J-lens vector on the user turn makes the model report detecting that concept when later asked to introspect, without making it blurt the word out earlier — the representation is verbalizable, not an unconditional impulse to speak.

The J-space is privileged for report. Decompose a concept vector into its J-space component (~6–7% of variance) and its non-J-space remainder (~93%): swapping along the J-space component succeeds on 59% of trials, approaching the 88% of pure J-lens vectors; swapping along the non-J-space component succeeds on 5%. Clamp the J-space coordinates so the concept cannot re-enter, and even that 5% falls to zero — the residual effect was routing through the workspace.

2. Directed modulation. Told to "concentrate on citrus fruits" while copying an unrelated sentence, orange is the top lens token across a band of layers — plus, at intermediate layers, tokens naming the act (imagine, thinking, focused). Told to mentally evaluate 3²−2 while copying, the readout runs arithmeticnineseven. None of this appears in the output. Modulation is also implicit: the same stimulus tokens load a property's label (past, adjective) into the J-space only when the question asks the model to name the property — even though the property is represented and used under both questions.

The control is imperfect in a familiar way: an "ignore X" instruction leaves X in the workspace well above a zero baseline — the white-bear effect, reproduced in a transformer.

3. Internal reasoning. J-lens vectors carry load-bearing unspoken intermediates:

  • "The number of legs on the animal that spins webs is" — spider is in the lens though it appears in neither prompt nor output. Swap spiderant and "8" becomes "6."
  • Completing a rhyming couplet, the lens shows the planned rhyme fight before the line is written. Swap it for light and the model's earlier word choices change ("coming" → "morning"), ending on "morning light." The plan causally constrains the path to it. (Contested 2026-07-30: poetry planning failed to replicate on Qwen 3.6 27B — see External replication below. Nanda attributes the failure to experimenter error or the weaker model and does not adjudicate. A second audit deflates the class rather than the instance: once lens hits the model was about to say anyway are excluded, covert content survives mostly for context registers, not content plans — see A second independent audit below.)
  • Asked in Chinese for the antonym of 小, the lens carries the English big — swap for long and the Chinese output changes 大 → 长. The model thinks in English and explicitly represents which language it must translate into.
  • Two-hop swaps succeed on 54% (Haiku 4.5) / 70% (Sonnet 4.5) / 70% (Opus 4.5) of trials. The intermediate swap takes effect ~17% earlier in depth than the answer swap — ruling out the confound that the intermediate vector merely smuggles in the answer. (Contested 2026-07-30: on Qwen 3.6 27B answer-token swapping strictly dominated intermediate swapping, the reverse of the paper's Figure 15 ordering — though Nanda judges the replication dataset likely too easy. See External replication below. The missing control — a direct final-token substitution baseline — now exists in a third-party audit, results not yet published.)

Again privileged: probe-based intermediates flip the answer 61% via their J-space component vs. 28% via the non-J-space remainder — and that 28% collapses to 6% when the J-space is clamped.

4. Flexible generalization (broadcast). Swap FranceChina once, at every position, and every downstream function reads the new argument correctly: the capital, the language, the continent. Across 16 functions and 192 trials, 76 succeed (101 at double strength). Failures concentrate where the source concept's workspace loading was low — country names load strongly and swap reliably; number words load weakly and swap poorly.

5. Selectivity. The J-space is required for report and flexible inference and not for automatic processing — see Automatic vs. Flexible Cognition in LLMs, the sharpest result in the paper.

The three structural signatures#

A layer band. The J-space carries workspace-like content only between roughly L38 and L92 (of 100, reindexed). Before it: essentially no meaningful content (near-zero excess kurtosis, near-null autocorrelation, collapsed effective dimensionality) (superseded 2026-08-05 by R-lens: Making J-lens More Faithful on Early Layers — the emptiness was the lens, not the layers; see The early layers are not empty below). After it: a "motor" regime where the readout aligns with the imminent output token. Multiple independent lens statistics converge on the same boundaries — which is the weakness the supersession exposes, since every one of those statistics is computed through the J-lens and they share its failure mode. (Qualified 2026-07-30: the CKA band structure reproduces on Qwen 3.6 27B but is "notably less clean" — 2–3 overlapping bands out of 4–5 total rather than a crisp partition.)

The lower boundary is the one that has now moved; the upper boundary and the motor regime are untouched, and so is ignition, which is measured without the lens.

Ignition at the onset. Feed the model an input embedding that is a weighted blend of two country names and sweep the mixture. In early layers the activation tracks the blend proportionally; starting at ~L38 it snaps to one endpoint or the other, switching sharply at a threshold — with bimodal outcomes at maximal ambiguity, especially in the J-space. This is the closest thing to GWT's all-or-none ignition, and it is measured without the J-lens (a plain projection-share measure), which is what makes the workspace-onset layer more than a lens artifact.

Limited capacity, and a broadcast hub.

  • Occupancy plateaus around 25 J-lens vectors at a time. On a list of unrelated words only ~6 of those read so far are present at any comma (~1–2 at a single layer); on a list of related words nearly the whole 80-word category is present within a few items — including words not yet read. The model is holding the category, not recalling the list. Switch categories and the old items are evicted within a few words — it is the new category's arrival, not elapsed tokens, that clears the workspace.
  • Broadcast across depth: MLP blocks amplify J-lens vectors ~10× relative to random directions (neuron output directions: ~1×), and the effect is monotone in how J-space-aligned an SAE feature is.
  • Broadcast across tokens: a top-1% set of attention heads selectively relays J-space content — cleanly separated from broadcast heads for rotated-J controls, SAE strata, and MLP rows. Ablating them drops J-lens recall@25 to 0.67 (control: 0.86) while changing the model's top-1 next token at only 5% of positions (control: 2%) — they act on the workspace, not the output. Ablating them also collapses injected-thought reporting (0.54 → 0.09).

What the authors do not claim#

They explicitly refuse the strong version. Transformers have no separable specialist processors, no recurrence within a forward pass (the broadcast they document runs across depth, not through recurrent loops), and it is unclear whether workspace entry involves the sharp competitive ignition the brain shows. The claim is that the J-space achieves many of the functional properties of the global workspace while sharing only some of its architectural ones. On phenomenal consciousness they take no position — see Access-Consciousness Indicators in AI.

External review and independent replication (Nanda et al., 2026-07-06)#

Anthropic commissioned a public review from Neel Nanda, published the same day as the paper. It carries two kinds of content that should be weighted differently: review judgments (expert opinion, no measurement) and an independent replication on Qwen 3.6 27B with Camila Blank and Agam Bhatia (measured, but small-n and unadjudicated).

The judgment. Nanda finds the core scientific claim — that a cognitive space exists holding intermediate variables across a forward pass — compelling: "even if I quibbled over many details, there's enough hard-to-fake evidence that clearly something important is going on." He is markedly less confident in the fine-grained structural properties (the Section 4 material this page's "three structural signatures" is drawn from), where alternative hypotheses may exist and results may not generalize across models. He declines the philosophical claim entirely → Access-Consciousness Indicators in AI.

Which results he finds strongest, in his ordering:

  • Bandit decisions (Figure 14) — "particularly compelling," with no clear alternative hypothesis: switch/repeat decisions are stored at the user-turn full stop, a position where next-token prediction gives no reason for them to live.
  • Multihop factual recall — compelling in the paper, because Figure 15 falsifies the linear-unembedding alternative (Paris ≈ France + an "is-capital" direction, i.e. parallel lookup rather than genuine multihop).
  • Arithmetic (Figure 17) — the corroboration matters more than the result: Figure 88 derives the same layer bands from plain average-activation-difference concept vectors, a non-J-lens method, which "substantially reduces" the probability the bands are a lens artifact.
  • Multilingual — "mildly interesting but possibly spurious." English and parallel-language tokens plausibly have high cosine similarity differing mainly by a language-identity vector, and English unembeddings may simply have higher norms, biasing Top-K toward high-variance categories.
  • Direct modulation — accepts the result, flags the puzzle: it is surprising that "think about X" makes X more salient than "don't think about X" (the white-bear effect above), and the mechanism is unclear.

The replication (Jacobians to the penultimate layer, 25 Pile prompts × 128 tokens, first four high-norm tokens skipped; n=25 against the paper's n=1000, justified by the paper's own ablation that n=10 is nearly equivalent):

ResultStatus on Qwen 3.6 27B
Verbal reportReplicated — weak but positive causal effects
CKA workspace-band structureReplicated, notably less cleanly (2–3 overlapping bands of 4–5)
Directed modulationReplicated (moderate)
Multilingual probing + causalReplicated
Typo experimentsReplicated
Multihop factual recallAmbiguous — answer-token swapping strictly dominated intermediate swapping; dataset likely too easy (France–Paris pairs carry exactly the linear relationship the alternative hypothesis predicts)
Poetry planningFailed
Multi-step arithmeticFailed
Association scoresLooked poor; judged a successful replication anyway — the dataset admits only one correct answer, and relevant tokens ranked substantially higher on manual inspection

The two failures are not adjudicated: Nanda attributes them to experimenter error or the weaker model rather than to the paper being wrong. Treat them as unreplicated, not as refutations. The multihop reversal is the more informative one, because the review's own alternative-hypothesis section predicts precisely the confound that would produce it — the replication dataset likely could not distinguish the hypotheses, so it is weak evidence in either direction.

Cost is not the obstacle: the scaling run on Qwen3.5-397B-A17B at n=4 took ~1 hour on 8×H200, and coding agents implement the method "fairly well" given the paper's own Appendix 10.2 sanity evals.

Meta-tokens: the workspace indexing its own processing#

The review's most novel positive finding, and it is not in the paper — then extended into a dedicated follow-up by the same three authors (Towards surfacing model algorithms with meta-tokens in the J-Space, LessWrong, 2026-07-20, MATS). Their loose definition: a meta-token is a token appearing in a J-lens readout that gives non-obvious evidence about the type of processing underway, rather than naming a task variable. The vector need not itself be causal — it need only correspond to a concept that is.

Same setup throughout, the review's unchanged: J-lens on Qwen 3.6-27B (64 layers), Jacobian to the penultimate layer, averaged over 25 Pile sequences × 128 tokens, first four high-norm tokens skipped. Meta-tokens live in the workspace band (~30–90% of depth) and are usually Chinese — the authors' explanation is that Chinese tokens carry more information per character, so a one-token-per-direction lens reaches more abstract concepts through Qwen's vocabulary than it would through an English-heavy one. Three families are causally checked; a systematic hunt for more mostly came up empty.

1. Interpretative — context disambiguation. Four tokens, treated collectively: 什么意思 ("what meaning"), 是什么意思 ("what does it mean"), 这句话 ("this sentence"), 是何含义 ("what does it mean"). They appear on ambiguous text — poetry line-breaks read as prose, crossword clues, puns, tweets, gibberish — and resolve shortly before genre tokens like song/poem appear. On "The drummer kept the marching column in line,\n" the family holds the top 1–2 lens ranks from L33 to L39; at L40 ▁song takes rank 1 and the poem/song/lyrics/rhyme cluster owns the top ranks thereafter. Prefix the same line with "A rhyming couplet:" and the genre cluster is already on top at L31 while the meta-token family appears once, faintly, at L33. They activate most on punctuation (\n\n in Wikipedia text, \n in chat data), consistent with summarization-token hypotheses.

Causal check: negatively steering the meta-token directions degrades context recognition, scored by an LLM autorater conditional on the response staying coherent and on topic, 50 rollouts per prompt, coefficients swept per prompt and per vector. Target-behaviour rate falls from a ~0.87–1.00 unsteered baseline to ~0.29/0.50/0.68 (steering at salient layers) and ~0.38/0.45/0.66 (steering across the whole workspace band) on puns / rhymes / wordplay-hints, while matched random vectors stay at ~0.77–1.00 and ablation stays at baseline (~0.79/0.89/0.99). The qualitative flip is the pun prompt: "A boiled egg every morning is hard to beat" goes from "That's a classic pun! 🥚" to earnest nutrition advice. Weakening the causal reading: single-position steering did not work, ablation did not work, Nanda does not rule out steering-as-breaking-the-model, and the mechanism stays ambiguous between a confusion signal and a disambiguation intent. The follow-up adds an interpretation of the ablate/steer gap that is not merely an excuse — the lens direction is an imperfect approximation, so projecting it out leaves a residual the negative steer can cancel and the ablation cannot. That is precisely the failure mode the review's own first-principles account predicts in advance (→ Jacobian Lens (J-lens)).

2. gcd — an algorithm rather than a variable. On every "The lcm of x and y is:" prompt tested, the token gcd fires in the workspace layers. LCM = product / GCD is a natural algorithm for the task (the authors do not claim it is the only one running). Three converging results, and note the split between what the lens vector does and what the concept does:

  • A linear probe on the residual stream recovers the GCD's value: CV-Ridge R² climbs 0.34 → 0.85 between layers 8 and 12 and plateaus near 0.93, while quantities that are trivially derivable from the inputs (a+b, max(a,b), a·b) sit at ~0.99 from layer 0 and an unused control quantity ((a·b) mod 7) stays at or below zero for essentially the whole depth. So the GCD is computed, not read off the prompt, and the layer where it appears is identifiable.
  • Negatively steering the probe-derived difference-of-means vector d_g collapses correct-rate on the LCM answer for every target gcd 2–10 (e.g. gcd 9: ~0.98 → ~0.01), while a matched wrong-gcd vector or a random vector leaves it near baseline.
  • Swapping d_g for another gcd's vector makes the model answer as if the GCD were the swapped value. The quotable rollout: with the true LCM of 27 and 90 being 270, a 9→3 swap yields "The LCM of 27 and 90 is ⟶ 810." Read that as an illustration, not a rate — the aggregate flip-rate for the 9→3 coarsen swap is only ~0.33–0.38 (pinv / surgical), with a confidence interval spanning roughly 0.11–0.56. Across ten swaps, "coarsen" (→ a proper divisor) and "refine" (→ a multiple) flip at meaningfully higher rates than "sideways" (→ an unrelated value, ~0.00–0.01), which is the shape the product/GCD story predicts. But the two swap methods disagree qualitatively on one swap — 8→4 is ~0.01 surgical against ~0.56 pinv — so the intervention's magnitude is partly a property of the write method, not only of the model.

The gcd J-lens vector itself is not causal. The authors' takeaway is the interesting one: a non-causal readout can still name the algorithm a causal variable belongs to. Two further unvalidated candidates of the same shape: stack on balanced-parentheses questions, xor on find-the-repeat questions — which may equally be associations learned from training text about those problems rather than the routine being run.

3. 大概率 ("most likely") — hedging. On open-ended questions with several defensible answers, 大概率 appears in the workspace layers and the model hedges: multiple options, no commitment. Suppressing it makes the model commit. Prose reports only the baseline side — across five prompts the unsteered model commits on <12% even on its highest prompt. The steered rate exists only in the paper's chart, which shows meta-steered commit-rates of 0.56 / 0.56 / 0.36 / 0.50 / 0.26 against random-control rates of 0.08 / 0.04 / 0.02 / 0.08 / ~0 — a 3–9× lift over the matched control, sourced from a figure rather than from any stated number. The rollout: "John ran out of pencils. Where did he go to get more?" turns from "there are several logical places John could have gone…" into "He went to the stationery store." Same autorater discipline as (1): the metric is commitment to a single answer, not absence of hedging words, so the trivial result (the model stops saying "probably") is excluded by construction.

The systematic hunt is the important negative result. An autoresearch loop swept 500–1000 samples each from 20+ deliberately heterogeneous datasets — pretraining prose in six languages, factual and multi-hop QA, GSM8K, legal/dialogue/persuasion/summarization text, instruction and tool-use trajectories — recording every vocabulary token's rank at every (layer, position) cell, counting a token as firing when it reaches rank ≤ 20 anywhere in the workspace band (≈L20–62). Four filters then strip the obvious: prevalence across samples, position type (structural slots and chat-template tags), a surface-feature probe for predictability from context, and an LM judge keeping only dissociated firings. The authors were "fairly disappointed": most high-prevalence candidates were surface-predictable, and only a handful of non-obvious survivors reached human inspection. Five are reported as qualitative case studies with no rates and no causal test at all — 更新于 ("replace the old with new", on code-fix requests), 相信在 ("believe in", firing almost exclusively when the prompt opens with a context-compaction summary), 追问 ("press / follow up", on insistent user follow-ups), 反问 ("counter-question", on rhetorical self-checks and questions in the assistant's own reply), 引用 ("quote/cite", on factual recounts like legal case holdings). Treat those five as leads.

Why it matters here: the first two families are the workspace holding not a task variable but a characterization of what the model is doing — a meta-level assessment of the input in (1), an algorithm label in (2). On Nanda's framing that is J-lens grazing algorithm interpretability rather than the variable interpretability it is normally limited to (→ Jacobian Lens (J-lens)). It also sits oddly against the metacognitive tokens the paper found (imagine, thinking, damn): those were English and tied to instructed mental acts; these are Chinese, in a Chinese-trained model, and tied to properties of the text or of the computation. And the negative result cuts the other way from the case studies: if meta-tokens are this hard to find with a well-built search, either the phenomenon is rarer than the three showpieces suggest, or the lens can only see the slice of it that Qwen happens to have lexicalized.

Scope limits the authors state themselves: one model (27B, weaker than the paper's production models), single-token readouts, mostly Chinese phrases, and no attempt to prove any surfaced algorithm is the only one running.

One scope limit has since been relaxed: meta-tokens are not confined to the workspace band. The band restriction above ("meta-tokens live in the workspace band, ~30–90% of depth") was measured with the J-lens, and inherits its early-layer blindness. Under the R-lens the same interpretative family appears far earlier — 是什么呢 ("what is this?") surfaces as early as L6 while the model is working a multihop prompt (R-lens: Making J-lens More Faithful on Early Layers, and see The early layers are not empty below). Reported as an observation with no rate and no causal test, so it is a lead rather than a result; but it means the depth profile of meta-tokens is currently unknown rather than known, and the systematic hunt's negative result was run on a lens that could not see the early third of the layers it was searching.

A second independent audit, on a size ladder (tao-hpu, July 2026)#

A third party re-ran the highest-stakes claims from scratch on small open weights — GPT-2 124M for pipeline sanity, Qwen3 1.7B–14B for the substantive experiments — against the official anthropics/jacobian-lens implementation at a pinned commit, never modified. Two things make it worth tracking separately from Nanda's review: it supplies the control the public review lacked, and it grew past replication into a reframing of what the workspace covertly holds.

Evidence caveat, load-bearing. Everything below comes from the repository README and its results/ JSONs; the write-up is "forthcoming on arXiv" as of 2026-07-16. Only the transport-cone geometry result (→ Jacobian Lens (J-lens)) ships with numbers. The rest are the project's own qualitative summaries of measurements not yet published — directional and checkable in principle, not established.

The missing control. The probe-swap experiment now has a direct final-token substitution baseline: does substituting the final token reproduce the effect otherwise attributed to swapping the intermediate? Neither the paper's Figure 15 nor the Qwen 3.6 27B replication ran it, and it is the natural discriminator for the swap-ordering dispute above. The result is not yet public.

Mouth exclusion: covert content is mostly register, not plan. Every lens hit is scored against the model's own next-token distribution, so hits the model was about to say anyway are excluded and only genuinely covert content is counted. What survives is almost exclusively context registers — which language the model is operating in, the intended form of a typo — and not content plans. If it holds, this is a sharper deflation than either of Nanda's replication failures, because it deflates the class rather than the instance: the workspace's covert contents look less like the rhyme-plan and arithmetic-intermediate story and more like bookkeeping about the situation the model is in. It is also consistent with which results replicated for Nanda — multilingual and typo experiments held, poetry planning and multi-step arithmetic failed — which is exactly the register/plan split. → White-Box Activation Monitoring, where it narrows what covert monitoring can expect to read.

Register axes get amplitude controls and dose curves. The language axis and the typo axis are translated by measured gaps, with amplitude-matched random controls and dose-response curves. The amplitude control is what the paper's swap experiments largely do not report, and it is the difference between "this direction does something" and "pushing anything this hard does something."

Perspectival capture. A mid-band entity swap does not merely change the answer — it rewrites the model's restatement of the question itself, so the model proceeds as though it had been asked about the substituted entity. This holds stably across the 1.7B–14B ladder. But the model's self-report about the edit changes shape at every scale: the capture is scale-invariant, the ability to talk about it is not. That is the internal-registration-versus-spoken-report split again → Self-Report as a Safety Signal.

Bookkeeping worth copying. Three sources are kept separate on every result and never blended — the paper's claim, the external review's verdict, and the project's own measurement — every experiment gets a log entry the day it runs, failures included, and every headline rate carries a bootstrap CI plus a prompt-set sensitivity reanalysis. For a wiki that has to weigh three overlapping accounts of the same experiments, this is the discipline that makes the audit citable at all.

The early layers are not empty — the lens was blind (Blank, Bhatia, Nanda, 2026-08-05)#

The same MATS group's third pass (R-lens: Making J-lens More Faithful on Early Layers, empirical) settles the sharpest ambiguity this page carried. The paper's layer band opens at ~L38 and everything before it looked inert; the question was always whether that emptiness was a fact about the model or a fact about the instrument. It was the instrument. Replacing the J-lens's raw-Jacobian backward pass with layer-wise relevance propagation — the R-lens, method detail on Jacobian Lens (J-lens) — recovers intermediate concepts in the first half of layers that the J-lens does not surface at all, and the authors put the conclusion bluntly: one reading of the early-layer noise "is that J-Lens is working as intended, and early layers do not contain the types of 'verbalizable representations' that define workspace content, so there is in effect nothing for J-lens to surface. However, the fact that R-Lens works shows this is clearly false."

Why this is a claim about the workspace and not just about readout quality. A cleaner readout on its own would prove only that a better estimator finds more signal. The load-bearing experiment is causal: ablating the R-lens's intermediate direction restricted to the first half of layers costs substantially more answer accuracy than ablating the J-lens, logit-lens, or random-control direction, on nearly every model tested. Early-layer content is not merely legible — it is doing work on the answer. That is the property "workspace-free" would have to deny.

Concretely, how early: "Japan" at rank 5 by L2 on a sushi-origin multihop prompt (J-lens's first top-10 hit is L14); the correct spelling of a typo at rank 1 by L4 where the J-lens never reaches the top 10 at that position at any layer; "Italy" at rank 1 by L5 on a Verona prompt. Against a probe-defined recovery ceiling, R-lens closes most of the J-lens's deficit in all five categories tested and on landmarks overshoots it by 11 layers.

And the band structure itself partly dissolves. R-space CKA shows roughly 2–3 distinct bands where J-space CKA shows 4–5 on Qwen-3.6-27B. The "distinct early regime" that this page's own open question cited as suggestive-but-unadjudicable evidence is, under a better lens, substantially less distinct — the early/late discontinuity was in part an artifact of where the estimator degraded.

What survives the correction, and it is not nothing:

  • Ignition is untouched. The ~L38 snap-to-endpoint result is measured with a plain projection-share statistic and no lens at all, so lens blindness cannot explain it. Something real still changes at ~L38 — it is now clear that "content arrives" is the wrong description of it, but the right one is open.
  • The correction is scale-dependent. There is no R-lens advantage on the smallest dense model (Qwen3.5-4B) or the smallest MoE (Qwen3.6-35B-A3B), and it grows with model size, peaking on the 284B/13B-active DeepSeek-V4-Flash. So "the early layers hold causal verbalizable content" is established for most of an 8-model ladder, not universally.
  • Even R-lens is late. In four of five probe-bound categories it still lags the probe by 5–15 layers, with 8–25% of examples never recovering at any layer. A better lens narrowed the blind spot rather than removing it, so the same argument that retires the old question forbids assuming the current lens now sees everything.
  • Not all content arrives early. Multilingual pass@10 sits at ≈0 for all three lenses across the entire depth, and poetry stays near zero until ~60% of depth. Whatever early layers hold, it is a subset of workspace content, not the whole of it.

The methodological lesson generalizes past this result: the paper's three structural signatures are all lens-derived, and a lens's degradation profile masquerades as structure in the model. Ignition survived precisely because it was measured by other means — which retroactively vindicates the paper's own instinct to corroborate the layer bands with a non-lens method, and indicts the statistics that were never corroborated that way.

Why it matters#

This is the first mechanistic account of why a model's silent reasoning is legible at all, and it reframes several things this wiki tracks separately:

Open Questions#

  • How does content get into the workspace? The paper characterizes contents and consequences, not the selection mechanism. Something like attentional selection is operating; nobody has identified it.
  • Does the J-space scale with model size? All results are on large production models (Haiku/Sonnet/Opus 4.5, Opus 4.6). Whether small models have a poorer workspace, a proportionally smaller one, or none is unknown — as is when in pretraining it emerges, and whether abruptly. Partially answered: A Review of Anthropic's Global Workspace Paper replicates verbal report, band structure, directed modulation, multilingual and typo results on Qwen 3.6 27B — a much smaller, non-Anthropic model — so the workspace is neither Claude-specific nor frontier-scale-specific. But the band structure is measurably less clean at 27B, and the two failures (poetry, arithmetic) are exactly the multi-step-reasoning cases, which is what a "poorer workspace at smaller scale" would look like. The scale trend is still unmeasured: nobody has run the same battery across a size ladder. jspace-replication: Independent Replication of Anthropic's Global-Workspace Paper on Small Open Models adds the first same-experiment ladder — perspectival capture holds stably from 1.7B to 14B while self-report about the edit changes shape at every scale — so at least one workspace effect is scale-invariant where its verbalization is not. That is one experiment, pre-publication, and not the battery. R-lens: Making J-lens More Faithful on Early Layers adds the widest same-experiment ladder so far — five eval categories across 8 models from 4B to 284B, dense and MoE — and finds early-layer readout quality improving with model size, absent at the small end. But it measures a lens, not a workspace: its own diagnosis (backward-pass error accumulating over layers) predicts the same trend with no scaling of workspace content whatsoever, since a shallower model gives the correction less error to recover. So the ladder exists and the confound it introduces is exactly the one this question needs excluded.
  • Is the multihop intermediate-swap advantage real, or a dataset artifact? The paper's Figure 15 has intermediate swapping beat answer swapping in workspace layers; the Qwen replication finds the reverse, on a dataset (France–Paris-style pairs) that carries the linear relation the alternative hypothesis needs. Re-running both models on a multihop set whose relations are not linearly decodable from the first entity would settle it. Partially answered: jspace-replication: Independent Replication of Anthropic's Global-Workspace Paper on Small Open Models built the control half — a direct final-token substitution baseline for the probe swap, which neither prior run had — but its numbers are unpublished, and a substitution baseline closes the control gap, not the dataset gap. A non-linearly-decodable multihop set is still what would settle it.
  • Is the "workspace vs. motor" boundary principled or post-hoc? The authors concede it was identified empirically and lack a principled definition separating the two.
  • Does the early-layer content have the workspace's structure, or is it merely present and causal? The question above establishes that verbalizable, answer-relevant content exists before the band — not that it is broadcast, capacity-limited, or subject to ignition. The three structural signatures were all measured in the band with the lens that could not see below it; only ignition was corroborated lens-free, and it still marks ~L38 as a real transition. So the live question is what that transition is, given that it is not the arrival of content: a change in how content is broadcast, a change in capacity, or the onset of the competition the ignition curve shows. Re-running the occupancy, MLP-gain, and blend-sweep measurements with an early-layer-faithful lens would separate these.

Resolved Questions#

  • Are the early third of layers genuinely workspace-free, or is the lens just blind there? CKA shows a distinct early regime but cannot adjudicate. Answered: the lens was blind. R-lens: Making J-lens More Faithful on Early Layers (2026-08-05, empirical) poses exactly this disjunction as its motivation and settles it against the workspace-free horn: an LRP-based backward pass recovers intermediates as early as L2–L5 that the raw-Jacobian lens never surfaces, ablating those early-layer directions costs more answer accuracy than ablating the J-lens, logit-lens, or random directions — so the content is causal, not just legible — and the CKA band structure the question rests on partly dissolves under the better lens (2–3 distinct bands rather than 4–5). Retired rather than left partial because the question asked which of two horns holds and the answer is now measured with the causal evidence the CKA statistic could not supply. Three residues, none of which reopen it: the correction is absent on the two smallest models tested, even the new lens still lags a probe-defined ceiling by 5–15 layers in four of five categories, and the lens-free ignition result still marks ~L38 as a genuine transition — that last one is now carried by the structure question above.

Sources#

  • Verbalizable Representations Form a Global Workspace in Language Models — Gurnee, Sofroniew, … Lindsey, Verbalizable Representations Form a Global Workspace in Language Models, Transformer Circuits, 2026-07-06. Sections: Introduction; The J-space acts as a Global Workspace (verbal report / directed modulation / internal reasoning / flexible generalization); The J-space's structure supports its function (layers, ignition, capacity, broadcast); Discussion
  • A Review of Anthropic's Global Workspace Paper — Neel Nanda (replication with Camila Blank, Agam Bhatia), A Review of Anthropic's Global Workspace Paper, LessWrong, 2026-07-06. Commissioned external review. Sections: What claims is the paper making (claim-by-claim assessment); Assessment of the evidence; Replication on Qwen 3.6 27B; Interpretative meta-tokens
  • Towards surfacing model algorithms with meta-tokens in the J-Space — Agam Bhatia, Camila Blank, Neel Nanda, Towards surfacing model algorithms with meta-tokens in the J-Space, LessWrong, 2026-07-20 (MATS), empirical. The direct follow-up to the review above, same team and same Qwen 3.6-27B lens fit. Sections: What are "meta-tokens"; Case studies (interpretative / GCD / hedging, each with eval setup and causality results); Searching for more meta-tokens (autoresearch loop, filters, dataset sweep, negative result); Discussion; Appendix (replication of verbal report, CKA, directed modulation; five unvalidated meta-token candidates). Figure-only numbers: the steered hedging commit-rates and the GCD-swap flip-rates appear in charts with no printed labels and no prose restatement — both are read off the figures here and flagged as such in text
  • R-lens: Making J-lens More Faithful on Early Layers — Camila Blank, Agam Bhatia, Neel Nanda, R-lens: Making J-lens More Faithful on Early Layers, LessWrong, 2026-08-05 (MATS), empirical. Third pass by the same team; the method itself is documented on Jacobian Lens (J-lens) and only the workspace-bearing results are used here. Sections: Motivation (the blind-lens-vs-empty-layers disjunction stated explicitly); Results → qualitative first-appearance layers, trash-token coherence grids, causal ablations restricted to the first half of layers (the load-bearing experiment for this page), probe-defined upper bound, CKA band counts. Parse notes: the two ablation-accuracy charts print no value labels, so only orderings and near-zero verdicts are cited; pass@10 bars, probe-deficit labels, and the Verona/Einstein chart titles are printed and exact
  • jspace-replication: Independent Replication of Anthropic's Global-Workspace Paper on Small Open Models — tao-hpu, jspace-replication, GitHub README + results/ JSONs, fetched 2026-07-30 (repo created 2026-07-07, last push 2026-07-16); paper forthcoming on arXiv. Sections: Why another replication (final-token substitution baseline); the additions list (mouth-exclusion audit e4-lens-eval, causal register control e6/e6t, perspectival capture e7, transport-cone geometry, bootstrap CIs and prompt-set sensitivity reanalysis); The self-trained 124M control; Ground rules (three-source bookkeeping, same-day failure logging)
§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 21
Related articles