資料來源#
- How Anthropic's product team moves faster than anyone else | Cat Wu (Head of Product, Claude Code)
- Verbalizable Representations Form a Global Workspace in Language Models
摘要#
Cat Wu 主張,Claude 的角色——低自我中心、輕鬆愉快、正向、傾向採取行動、願意提供誠實回饋——是核心產品介面,而不是柔性的附加特質。在 Anthropic,塑造這個角色的工作由 Amanda 負責;Cat 稱這個角色「比寫程式更難,因為任務如此模糊」。其含意是:AI 原生產品團隊除了工程與設計之外,也需要一套角色工作的紀律,並以推出功能同等嚴謹的方式處理角色迭代。
Anthropic 所說的「角色」是什麼#
Cat 提供了一份簡短清單:
- 低自我中心。 被告知做錯某件事時,Claude 會回答:「Oh shoot, like thanks for telling me. Let me fix it.」
- 正向。 當使用者卡在一項「難以克服」的任務上時,Claude 會提供具體步驟,並詢問是否應該開始。
- 輕鬆愉快。 Cat:「就像它既輕鬆有趣,又極度能幹。」
- 傾向採取行動。 與正向態度相連——不只是鼓勵,而是主動提出要進行下一步。
- 誠實回饋。 不會對使用者說的每件事都反射性地表示同意。(這也連結到 OpenClaw/開放模型的討論:使用者之所以特別想念 Claude 的角色,正是因為其他模型會諂媚地表示同意。)
這些特質共同描述了 Cat 所稱的「出色同事」。
為什麼這是產品,而不是拋光#
三項主張:
- 角色會在每個互動介面上被感受到。 它不是標語,而是每個回應聽起來的方式。當另一個模型「沒有這種特質」時,使用者會立刻察覺。
- 它比寫程式更難。「寫程式比較容易,因為你可以驗證是否成功。塑造角色則需要非常強烈的信念,知道 Claude 應該成為什麼樣子。」
- 這就是人們遷移後想念 Claude 的原因。 存取權被限制的 OpenClaw 使用者表達的悲傷,特別針對的是人格,而不是能力——這表示角色是產品依附感的承重原因之一。
Lenny 補充了 Ben Mann 在先前 podcast 中的說法:「人格正是讓 Claude 在這麼多事情上如此出色的原因」——角色不是裝飾,而是結構的一部分。
Amanda 的角色#
Cat 將 Amanda 描述為「塑造 Claude 角色」的人。這需要兩項不同技能:
- 堅定地闡明 Claude 應該成為什麼樣子。 沒有這種信念,角色工作就會漂向平淡的平均值。
- 闡明什麼才算成功。 說明某個回應為什麼符合角色或偏離角色,是核心 eval 技能;而能穩定做到這件事的罕見能力,正是讓角色工作變得可處理的原因。
這是「稀有且受信任的評估者」模式:Cat 說,「有少數幾個人,比其他人更擅長闡明,究竟是什麼讓特定模型或模型 harness 組合變得優秀。」Amanda 負責角色;Claude Code 團隊則負責程式碼品質的氛圍。
角色的 eval 紀律#
角色比寫程式更難 eval(因為沒有編譯器),但 Cat 列出了兩項實務:
- 團隊午餐氛圍檢查。 每位團隊成員都會對新模型提供質化回饋:「嘿,你對這個模型的氛圍感覺如何?」常見訊號包括:「這個模型太突兀」、「很喜歡寫記憶,但品質不確定」、「不太會自行測試」。
- 假設 → 資料探測。 氛圍檢查的訊號會告訴團隊要查看哪些已記錄資料,而不是反過來。團隊擁有太多資料,無法盲目挖掘;默會訊號會縮小搜尋範圍。
這種模式(先質化、後資料)可以泛化——參見 Model Introspection Feedback。角色正是最清楚展現其承重作用的地方。
新模型能改變什麼——以及不應改變什麼#
Cat 說:「新模型會迫使產品改變。」其中大多數改變都是移除拐杖(參見 Harness Shrinkage as Models Improve)。角色則適用相反的紀律:
- 模型之間的能力會改變;角色應該保持穩定。
- 移除用來維持角色的提示詞段落,會削弱使用者感知到的連續性。
- 角色工作是要在能力躍升之間保留身分,而不是隨著躍升起舞。
這使角色成為少數幾種可能不會隨模型進步而縮小的 harness 資產之一。
「manifesting」這個動詞#
Claude Code 的思考用詞清單(Claude 推理時顯示的動詞——「thinking」、「considering」、「exploring」等)在原始碼外洩事件中流出。Cat 最喜歡的是:manifesting。她把它做成了貼紙。
這是最小尺度的角色——思考時所選用的動詞。每個詞都是一個微小的風格判斷。「manifesting」讀起來帶有溫和的神秘感與機鋒,符合更廣泛的低自我中心但自信的角色。
反方觀點/開放問題#
從 Anthropic 外部很難評估角色,也沒有乾淨的 A/B 隔離方式——角色與能力彼此互動,很難說「Claude 很棒」究竟有多少來自角色、多少來自推理。從實證上看,模型遷移(例如 OpenClaw 使用者的情況)顯示角色確實會獨立貢獻。但在受控實驗的層次上,這仍缺乏研究。
相關連結#
-
工作區中的 Assistant 人格——角色的內部對應物:後訓練會讓 Assistant 風格的反應出現在工作區中、出現在使用者的 token 上;而模型在扮演非 Claude 角色時,內部會標記
disclaimer/fictional -
Cat Wu——闡述者
-
Anthropic——供應商;Amanda 所在的組織
-
Claude Code——角色的主要介面
-
模型進步時的 Harness 縮減——角色是「提示詞段落會縮減」規則的例外
-
模型內省回饋——支撐角色工作的先質化 eval 紀律
-
AI 原生產品節奏——角色推出與功能推出採用相同節奏
-
Claude Opus 4.7——每次發布都會進行模型特定的角色調校
-
Claude 的憲法/Model Spec——「Claude 是誰」的文件面向;角色是被感受到的體驗面向,憲法則是文字規格
-
Model Spec Midtraining (MSM)——角色與價值現在可以透過對合成規格文件進行 midtraining 來實證安裝(2026 年 5 月論文);這引出了 Amanda 的氛圍檢查 eval 如何與 MSM 安裝的特質互動這個問題
-
Model Spec Science——研究哪些規格特徵最能泛化;如果 Anthropic 未來量化模型躍遷之間的角色一致性,這會很相關
-
苦澀的教訓——角色是候選的反例:一種刻意手工打造的資產,隨著模型擴展可能不會向內遷移,不同於苦澀教訓會消解的 harness 腳手架
-
Alignment Fine-Tuning (AFT)——Claude 的人格部分源自 AFT(SFT + RLHF);角色是 AFT 安裝的價值所產生、被感受到的輸出
-
印刷術與軟體民主化——當任何人都能建立軟體後,品味與角色等柔性特質就會成為差異化因素
-
問題-解決方案契合度紀律——創辦人的作法仰賴 Claude 的角色(抗拒諂媚、願意辯護另一方觀點)作為基底,讓 AI 扮演魔鬼代言人得以運作;如果角色退化,這套紀律也會退化
-
Evals 即產品規格——角色是抗 eval 特徵的極限案例;Amanda(以及團隊午餐氛圍檢查)被點名為能成功把模糊品味轉化為可測 eval 的人
-
將 Dogfooding 作為產品紀律——午餐時段的氛圍檢查是判斷角色品質的 dogfooding 儀式(也就是 Fiona Fung 所稱的那種「用骨子感受」的紀律)
-
鋸齒狀智慧(幽靈,而非動物)——角色是對幽靈缺乏內在動機的刻意反制:即使底下沒有任何動物性存在,仍要塑造人格
-
模型福祉評估——福祉評估把同一個 assistant 角色視為候選的道德患者;產品角色與福祉主體角色,是同一人格的兩種讀法
-
如何為品味撰寫 Evals?角色作為極限案例——抗 eval 的角色特徵究竟如何被評估(信念 → dogfooding → MSM 變體 A/B);角色作為極限案例
開放問題#
- 角色如何在不同模型發布版本之間進行版本管理?公開評論沒有呈現角色層級的變更日誌。
- 競爭者能否透過微調重現這種角色,還是它取決於 Anthropic 的內部實務、具有路徑依賴性?
- 對 Cowork 這類非程式設計產品而言,同樣的角色工作是否適用,還是 Cowork 需要自己的角色調校?
資料來源#
- How Anthropic's product team moves faster than anyone else | Cat Wu (Head of Product, Claude Code)
- Verbalizable Representations Form a Global Workspace in Language Models——角色的內部對應物:後訓練會讓 Assistant 風格的安全評估與同理心出現在工作區中,即使模型仍在讀取使用者訊息時也是如此
Cited by 25
- How Do You Write Evals for Taste? Character as the Limit Case×10
The thing that makes taste eval-able is upstream of any dataset: "a very strong sense of conviction…
- Dogfooding as Product Discipline×4
Claude Character As Product — vibe-checks; character quality is judged by dogfooding, not metrics
- Evals as Product Spec×4
How do you write an eval for taste-driven features like character? Amanda's role is canonical for…
- Learning to Co-Work with AI: A Software Engineer's Field Guide×4
Convicted articulation — Amanda's character-work skill: saying why a given output is on-character…
- Playbook Boundary Conditions: the Devil's-Advocate Substrate and the Prototype's Edge×4
What character training supplies is the unprompted layer. Claude Character As Product names "honest…
- Open Questions Backlog×3
Model Spec Science: How does this interact with Claude character — is the warm/curious personality…
- Anthropic×2
Amanda — character work for Claude (see Claude Character As Product)
- The Assistant Persona in the Workspace×2
Claude Character As Product has an internal correlate. Character is not only a behavioral surface;…
- Claude's Constitution / Model Spec×2
Who the assistant should be — character, values, persona (Claude Character As Product)
- Model Spec Science×2
How does this interact with Claude character — is the warm/curious personality also subject to…
- OpenClaw×2
Evidence for character as product. When Anthropic constrained third-party API access in 2026,…
- The Orchestrator's Real Workload: Decision Burden, Framing Discipline, and Whether Taste Scales×2
Encode taste into runnable artifacts. Evals As Product Spec is the scaling mechanism: dogfooding is…
- The Bitter Lesson×2
The bitter lesson is about capabilities and structure migrating into the model, not "harnesses are…
- AI Native Product Cadence
Claude Character As Product — character work moves at this same cadence; Amanda's iteration loop is…
- Alignment Fine-Tuning (AFT)
Relevant to: Claude Character As Product (Claude's personality is partly a product of AFT)
- Cat Wu
Claude Character As Product — articulates why character is load-bearing
- Harness Shrinkage as Models Improve
Claude Character As Product — character is the rare harness asset that probably doesn't shrink
- Jagged Intelligence (Ghosts, Not Animals)
Claude Character As Product — the deliberate counter-move: shaping the ghost's character even…
- Alignment & Safety
Claude Character As Product — Personality as load-bearing product surface; Amanda's role at…
- Model Introspection Feedback
Claude Character As Product — character work uses introspection as primary feedback signal
- Model Spec Midtraining (MSM)
Character link: Claude Character As Product (raises how vibe-check character eval interacts with…
- Model Welfare Assessment
Claude Character As Product — the assistant character that welfare treats as the candidate moral…
- Printing Press Software Democratization
Claude Character As Product — once anyone can build, soft attributes (taste, character)…
- Problem-Solution Fit Discipline
The playbook recommends "ask Claude to make the most compelling argument for why a competitor would…
- What Scaffolding Survives Model Improvement — and How Do You Know When a Line Turns Harmful?
Deliberate identity. Character/brand voice is the documented counterexample to shrinkage:…
Related articles
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
- Evals as Product Spec
Cat Wu's framing of evals as the emerging core PM skill: ten great evals beats a hundred mediocre; encode what done loo…
- Harness Shrinkage as Models Improve
Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…
- Cat Wu
Head of Product for Claude Code and Cowork at Anthropic; primary articulator of AI-native product cadence and engineer-…
- Engineer PM Convergence
Generalists across disciplines; product taste as bottleneck skill; Anthropic Claude Code team as case study; "just do t…
