H
Howardism
Plate IIEntities機器翻譯 · machine-translatedENHOWARDISM

Jack Lindsey

Anthropic 可解釋性研究員;全域工作空間論文的通訊作者、Jacobian Lens 的共同發起者,也是執行定向調節與訓練後差異比較實驗的人,這些實驗將一種讀出方法轉化為關於模型認知的主張

Article metadata
Publication details
Published:July 11, 2026
Filed:Entity
Domain:Entities
Tags:EntityPersonAnthropicInterpretability Researcher
Reading:2 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Jack Lindsey 的插圖

資料來源#

摘要#

人物。 Anthropic 可解釋性團隊的研究員,也是 Verbalizable Representations Form a Global Workspace in Language Models(Transformer Circuits,2026 年 7 月)的通訊作者。他與 Wes Gurnee 共同構思了 Jacobian Lens (J-lens),以及將可言說表徵與意識存取連結起來的猜想。

貢獻#

根據論文的作者貢獻章節:

  • 與 Wes Gurnee 共同構思 Jacobian Lens (J-lens) 方法,以及可言說性↔意識存取之間的關聯
  • 執行早期實驗,研究模型能否直接調節自己的 J-space(依照指示在心中保持某個概念),以及訓練後對透鏡讀出的影響——這些結果後來成為工作空間中的助理人格
  • 與 Nicholas Sofroniew 一同提出將 J-space 與全域工作空間理論連結的實驗——這一步將可解釋性讀出轉化為關於模型認知功能組織的主張

論文也引用了他先前關於 transcoder 與歸因的研究(透過 J-lens 重新檢視的算術特徵來自 Lindsey 等人的研究)。

相關連結#

資料來源#

§ end
Cited by 4
Related articles