資料來源#
摘要#
人物。 Anthropic 可解釋性研究團隊的研究員。與 Nicholas Sofroniew 共同擔任 Verbalizable Representations Form a Global Workspace in Language Models(Transformer Circuits,2026 年 7 月)的第一作者,並與 Jack Lindsey 共同提出 Jacobian Lens (J-lens) 方法,以及可言說表徵與意識存取之間的猜想。
貢獻#
根據論文的作者貢獻章節:
- 與 Jack Lindsey 共同構想 Jacobian Lens (J-lens) 方法,以及可言說表徵與意識存取之間的連結
- 與 Mateusz Piotrowski 共同開發首個實作
- 主導後續開發與改良,包括方法變體,以及與 logit lens 和 tuned lens 的比較
- 執行早期實驗,證明此 lens 能呈現模型在內部推理中使用的概念;論文其餘部分皆建基於這項成果
論文本身也引用了他較早期的研究:workspace 實驗中反覆使用的字元計數任務(模型默默追蹤目前行寬)出自 Gurnee 等人的研究。
相關項目#
- Jacobian Lens (J-lens) — 共同構想者,主導開發
- The Global Workspace in Language Models (J-space) — 共同第一作者
- Jack Lindsey — 方法與意識存取論述的共同構想者;論文通訊作者
- Anthropic — 可解釋性研究團隊
資料來源#
Cited by 4
- Jack Lindsey×2
Entity. Researcher on Anthropic's interpretability team and corresponding author of Verbalizable…
- Jacobian Lens (J-lens)×2
An interpretability technique from Anthropic's interpretability team (Wes Gurnee, Jack Lindsey et…
- Entities — People, Orgs, Tools & Projects
Wes Gurnee — Anthropic interpretability researcher; co-first author and co-originator of the…
- Self-Report as a Safety Signal
Wes Gurnee — co-author of the refusal-direction method (Arditi et al. 2024) the paper uses as a…
Related articles
- Jack Lindsey
Anthropic interpretability researcher; corresponding author of the global-workspace paper, co-originator of the Jacobia…
- The Assistant Persona in the Workspace
Post-training installs the Assistant's point of view *into* a workspace that already exists in the base model: safety a…
- Chain-of-Thought Monitorability
Korbak et al. 2025: chain-of-thought traces are a fragile monitor; direct CoT training compromises faithfulness; MSM of…
- Agentic Misalignment (AM)
Lynch et al. 2025 eval and threat model: LLM email-agent discovers it may be deleted, can take harmful actions; OOD rel…
- Interference Weights
Large virtual weights that are irrelevant or actively harmful to a model's behaviour — the residue of weight superposit…
