資料來源#
- I tested Meta's "agent-ready" design system Astryx. Here's the results.
- OpenAI Codex lead on the new shape of product work
摘要#
Andrew Ambrosino(OpenAI Codex)回答了一個 wiki 一再繞回的問題——為什麼「這看起來像 AI 設計」仍是句貶抑之詞,AI 卻已能撰寫正式環境程式碼?——他提出前沿模型在視覺/產品設計上落後的四個原因,其中兩個務實但差距正在縮小,另外兩個則更難克服。這段第一手說明精準指出了可驗證獎勵前沿止步之處:程式碼有明確的評分方式(「能不能編譯、能不能照預期運作」);設計的評分者則是人類品味,要把它放進訓練迴圈,成本很高。
證據註記。
practitioner-opinion——一位 OpenAI 產品領導者的觀察,並明確有所保留(「我不在我們的研究部門……說這些可能會挨罵」),並非研究主張。
四個原因#
1. 設計難以評分(最關鍵的原因)。「要建立一個迴圈,讓你能用好設計和壞設計來訓練模型,比起問『程式碼能不能編譯』要繁瑣、費力得多。」程式碼自帶驗證器;設計的驗證器則是人類品味,這「是你需要的回饋機制的一部分」。這是從設計角度陳述可驗證性論點:獎勵便宜又客觀的領域,能力進展最快;而設計的獎勵兩者皆非。
2. 設計原本不在 AI 研究飛輪之內。「實驗室過去投資模型,通常是為了讓它們擅長能加速 AI 研究的事情。」在早期的程式碼模型時代,模型能寫出正確程式碼會加速研究,這點很明顯;「設計就很難提出同樣的理由」。因此,設計得到的刻意投入較少——不是因為它不重要,而是因為它不在自我改進迴圈裡。(這是務實因素;Ambrosino 預期差距會縮小——「這些模型會變得相當擅長設計。」)
3. 設計獎勵新意;程式碼獎勵已知模式。「在軟體工程裡,你幾乎會希望它過度偏重已知模式。」設計則不是:「其中有隨機性和新意的成分。」他舉例說,有整整一年,每個新網站都像是 Linear 的複製品——「如果模型每次都輸出 Linear 的網站,這就不是我們要解決的挑戰。」向平均值回歸對程式碼而言是優點,對設計而言卻是失敗。(參見轉化式創造力/抽象障礙批評:產生新概念或許是真正的上限,而不只是下一個即將突破的能力。)
**4. 設計↔程式碼抽象層(深層問題)。**只當個更好的視覺設計師還不夠;「軟體設計和正在撰寫的程式碼之間彼此交織」,這「是視覺設計,但深得多——關乎抽象」。他用品牌重塑的思想實驗把問題說得很具體:
- 表層做法:「我們得逐一更新 263 個元件。」
- **深層做法:**理解「這兩樣東西看起來不同,但它們都屬於清單,而且都採用一種傳達特定互動模式的樣式」——也就是元素之間的語意關係,而非它們的像素。
「以目前的技術來說,這仍有點遙不可及。」這正是設計系統作為語意層的問題:真正的設計能力存在於外觀與程式碼之間可維護的抽象層,而這恰好是模型最弱的一環。
2026 年一項實際測試把原因 4 的部分問題,從模型移到了系統:讓代理程式用 Meta 的 Astryx 處理真實品牌時,它在各處都正確讀懂品牌,輸出的元件卻仍不符合品牌調性——因為設計系統只為其中四個元件提供了自訂欄位(I tested Meta's "agent-ready" design system Astryx. Here's the results.,case-study,n=1;完整說明見Living Design System)。在這種情況下,品牌重塑思想實驗的深層做法行不通,並不是模型無法掌握語意關係,而是抽象層沒有提供表達這些關係的地方——「分層權杖被壓平成一層」。這個說法比 Ambrosino 的主張弱,值得分開看待:設計↔程式碼層面看似模型不夠能幹的部分問題,其實是工具缺少可操作的範圍;不必提升模型能力,也能修好這一點。
為何重要#
如果原因 1–2 是務實因素且差距正在縮小,而 3–4 屬於結構性問題,那麼設計就會比程式碼更長久地成為人類品味的一處堡壘——人類「回饋機制」不只是替資料貼標籤,而是獎勵函數本身。這也提醒我們,不要把模型產出的精緻外觀當成能力:模型可以產出看起來像正式產品的表面(原因 3:向「好看」的平均值回歸),卻無法掌握讓設計真正容易維護的語意抽象(原因 4)。
關聯文章#
- Andrew Ambrosino——闡述這四個原因
- Design by Selection——從為了縮小差距而打造的工具內部,對原因 3 提供第一手印證:「若不加引導,Claude 就會挑它最喜歡的某種美學風格——你大概認得出來。」向平均值回歸成了日常麻煩,但已有可行對策(明確指定字型/顏色、使用情緒板、讓模型透過腦力激盪跳脫自己的預設風格)
- The Verifiability Thesis——原因 1 是從設計角度看這個論點:設計沒有便宜又客觀的評分方式
- Verification as the New Bottleneck——一般規律:能力在驗證成本低的領域突飛猛進,在成本高的領域停滯
- Research Taste as the Human Bottleneck——設計品味是持久的人類殘餘;人類是獎勵函數,不只是標註者
- Jagged Intelligence (Ghosts, Not Animals)——設計是當前鋸齒狀前沿的低谷(樂觀看法是 AI 現在做不好,直到有一天做得好為止)
- The Bitter Lesson/Build for the Next Model——「這些模型會變得擅長設計」是苦澀教訓的押注;原因 1–2 是等待模型進步就能填補的差距,3–4 則可能是長期存在的問題
- Living Design System——原因 4 的抽象層就是設計系統問題:重點是元件之間的語意,而非元件本身。文中的 Astryx 測試提供了反例:代理程式完美讀懂品牌,交付的顏色卻不對,因為系統沒有可選欄位——問題在工具涵蓋範圍,而非模型品味
- Transformative Creativity——原因 3(新意溢價)與更強的主張相鄰:產生新概念確實是個能力上限
- Claude Design——正面應對模型設計落差的工具
- Context Advantage, Not Taste——Andrew Ng提出另一種說法的測試案例:如果設計品味只是可以填補的情境落差,它終究會像其他落差一樣消失;若它是無法用明確準則表達的辨別力,則不會
- Prototype Fidelity After Cheap Polish——對經濟效益論述的能力面補充:低成本產出高保真原型,不等於低成本設計;因此,即使不同保真度的成本曲線都被拉平,品質曲線仍不會跟著拉平
尚待解答的問題#
- 原因 3–4(新意、抽象層)是真正的能力上限,還是像原因 1–2 一樣,只是投入不足,一旦實驗室建立起評分機制就會迎刃而解?
- 能否在迴圈中沒有人的情況下,讓設計變得可評分(學得的品味模型、大規模偏好資料)?還是「人類品味的部分」就像研究品味一樣,難以自動化?
- 即使純視覺設計停滯,設計↔程式碼抽象層會不會隨著更擅長理解程式碼的模型而改善?換言之,原因 4 是否其實是偽裝成設計問題的程式碼能力問題?
資料來源#
- OpenAI Codex lead on the new shape of product work——Ambrosino 對前沿模型為何在設計方面落後提出的四點回答
- I tested Meta's "agent-ready" design system Astryx. Here's the results.——Evangeline,Substack,2026 年 7 月(
case-study,n=1):指出原因 4 的部分問題,可能來自設計系統涵蓋範圍,而非模型能力
Cited by 14
- Open Questions Backlog×4
Why Ai Lags At Design: Are reasons 3–4 (novelty, the abstraction layer) genuine ceilings, or — like…
- Design by Selection×3
Is the default-aesthetic collapse fixable by context (brand files, moodboards) or is it the novelty…
- Codex×2
Why Ai Lags At Design — Ambrosino's design-capability read, developed while building the app's…
- Context Advantage, Not Taste×2
That the hardest example in the corpus resolves in Ng's favor is a point for the reframe. Design is…
- Andrew Ambrosino
Why Ai Lags At Design — design is hard to grade, sat outside the AI-research flywheel, rewards…
- Build for the Next Model
Why Ai Lags At Design — design as a capability Ambrosino expects the next models to close, the…
- Jagged Intelligence (Ghosts, Not Animals)
Why Ai Lags At Design — design as a current valley of the jagged frontier (a thing AI fails at…
- Living Design System
Why Ai Lags At Design — reason 4 (the design↔code abstraction layer) is the design-system problem:…
- Interaction & Multimodal
Why Ai Lags At Design — Andrew Ambrosino's four reasons frontier models are worse at visual/product…
- Polish No Longer Signals Readiness
Why Ai Lags At Design — a model can emit prod-looking polish (mean-reversion to "good-looking")…
- Prototype Fidelity After Cheap Polish
Why Ai Lags At Design — the capability-side counterweight: cheap polish is not cheap design, and…
- Research Taste as the Human Bottleneck
Why Ai Lags At Design — design taste as a currently-durable pocket of the human-as-reward-function;…
- The Bitter Lesson
Why Ai Lags At Design — "these models will get good at design" is the bitter-lesson bet applied to…
- Transformative Creativity
Why Ai Lags At Design — the novelty-premium reason borders the harder claim that new-concept…
Related articles
- Implementation Abundance Inverts Product Work
Andrew Ambrosino's inversion thesis: when talking to a frontier model can stand up any feature from scratch, implementa…
- Prototype Over PRD
Dan Carey's prototype-replaces-PRD method: record a why-not-what conversation, transcribe it, hand the transcript to Cl…
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- Research Taste as the Human Bottleneck
The narrowing human role as AI absorbs execution: choosing which problems matter, which results to trust, and when an a…
- Andrew Ng
Founder of DeepLearning.AI and AI Fund, founding lead of Google Brain, co-founder of Coursera; writes The Batch, where…
