資料來源#
- 3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse
- Models are worse at reviewing their own code
- Trust but Verify? Uncovering the Security Debt of Autonomous Coding Agents
這是什麼#
Greptile 銷售一款會對提取要求提出意見的 AI 程式碼審查代理程式。它不是模型實驗室:主要審查代理程式「通常是來自 OpenAI 或 Anthropic 的前沿模型」,因此產品本身是代理程式 harness、儲存庫脈絡與路由機制,而不是審查模型的權重。這種架構讓它得以進行研究——替換底層模型只需變更設定,而這正是其核心研究採用的操弄方式。
發表內容#
- 「Models are worse at reviewing their own code」(Rodrigo Caridad、研究團隊,2026-07-21)——兩組各含 500 個 PR 的真值資料集,一組由 Claude Code 撰寫,另一組由 Codex 撰寫,內含約 1,500 則已驗證的高嚴重性錯誤評論。每個前沿模型在審查自家系列模型撰寫的程式碼時,找出的錯誤都比審查另一系列模型撰寫的程式碼少:Opus 4.7 同模型為 53.7%,跨模型為 60.0%;GPT 5.5 則為 50.5% 對 62.0%。參見 同模型審查盲點。
- Model Inversion——根據這項發現打造的產品功能:從 PR 的軌跡辨識撰寫者所用的代理程式,並將審查工作路由給另一家供應商的模型。
- 一份 State of AI coding 報告,以及一系列解說 AI 程式碼審查的部落格文章,這些內容在實務工作者的討論中流傳(參見 3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse),也伴隨著反覆出現的「Greptile 替代方案?」討論串——這項工具的採用程度已足以引發引用,也帶來抱怨流失客戶的聲音。
Tran 等人的機密洩漏研究也將它列為測量對象,而非作者:Greptile 是曾對 74 組真實線上憑證中的任一組發表評論的七個機器人與人工審查者之一;這些對象合計的評論率構成了 代理程式生成程式碼的安全負債 中 18.9% 的下限。
如何解讀(證據立場)#
Greptile 的旗艦研究在編譯時被列為 case-study,由原始文件的 empirical 下修——該列持續顯示的 [evidence] lint 警告,正是這項修正,而非資料漂移。這是實際進行的研究,但真值資料由供應商自行建構,依據的是「情緒分析、按讚/倒讚比例,以及 git 考古」,卻沒有發布研究產物、標註流程、評分者間一致性、對負責配對的 LLM 評審器進行驗證,也沒有提供差距為 6 至 12 個百分點時的信賴區間。其中一組的提示詞還曾針對結果指標調整,直到召回率恢復。利益衝突也十分直接:跨供應商路由正是單一模型審查者無法提供的能力,因此這項發現也替據此打造的產品背書。受測對象是兩款第三方模型,分別透過各自供應商的審查功能執行;這就是其證據等級列為 case-study 而非 vendor-claim 的原因。
相信趨勢方向;具體數值則應歸因於 Greptile;能判別差異的複現研究——在兩組語料上測試第三個模型系列——仍是同模型審查盲點頁面上的一項待解問題。
相關連結#
- 同模型審查盲點——該研究測量並命名的發現,也是記錄其證據等級修正的頁面
- 審查作為控制點——其研究為自動審查者能力調節變數提供了一個系譜變數,而理論原先將此因素排除在外
- 最佳化器與評估器解耦——其資料首次以模型系譜而非角色,測試製作者/檢查者規則
- 代理程式生成程式碼的安全負債——其中將該機器人列為七個評論參與者之一,共同構成 18.9% 的偵測率下限
- Claude Opus 4.7、Codex——該研究所用語料分別出自這兩個撰寫代理程式
資料來源#
Cited by 3
- Deterministic Engineering for Agent Code Review
Anthropic, Openai, Greptile — the vendors whose shipped review features the corpus's three…
- Entities — People, Orgs, Tools & Projects
Greptile — AI code-review agent vendor that runs frontier models from OpenAI and Anthropic under…
- Same-Model Review Blindness
Rodrigo Caridad of Greptile's research team reports that a frontier model reviewing code its own…
Related articles
- Agent Review Comment Resolution
Cynthia et al. (Saskatchewan/SMU/Monash, arXiv 2607.21997): 54,713 agent review comments from Copilot, Cursor and Codex…
- Writer/Reviewer vs Agent-to-Agent Review
The two patterns are one architecture differing on venue, so the corpus's evidence attaches to four decomposed axes rat…
- Claude Code
Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…
- Risk-Tiered Auto-Approval
PostHog's StampHog: a merge-gate that auto-approves PRs passing four ordered checks (PR state, blast-radius deny-list,…
- Same-Model Review Blindness
Greptile's Rodrigo Caridad on two 500-PR labelled datasets (~1,500 verified high-severity bugs): each frontier model ca…
