資料來源#
摘要#
Perplexity 是一家 AI 答案引擎/搜尋公司。在本文語料中,它以 DRACO benchmark 的作者身分出現(arXiv:2602.11685,2026 年 2 月發布,並有一位 Harvard 共同作者),也是 Perplexity Deep Research 的開發者。這套代理式 深度研究 系統在該 benchmark 的每個領域與評分軸向都名列前茅。它是本 wiki 中第一家以自家部署流量打造源自生產環境的評估的供應商。
Perplexity 在本文語料中的功能#
- Perplexity Deep Research — 深度研究代理程式,會拆解查詢、反覆從多個來源擷取資料,並綜合成附有引文的報告。在 DRACO 上,其標準化分數為 70.5%(以 Opus 4.6 為基礎模型),通過率為 72.8%,領先 Gemini Deep Research、OpenAI Deep Research(o3 / o4-mini),以及未經編排、搭配工具的 Claude Opus 4.5/4.6。其特點是深度研究系統中分數最高、延遲最低,輸入 token 用量也最大(每項任務約 779k)——偏重資料擷取,輸出較精簡。
- DRACO — Perplexity 從自家 Deep Research 在 2025 年 9 月至 10 月的數千萬筆查詢中抽樣,接著去識別化、擴增、篩選並整理成 100 項由專家依評分規準評分的任務(源自生產環境的評估)。已在 Hugging Face 公開發布。
值得注意的結構性事實:既是客戶,也是競爭者#
Perplexity Deep Research 以 Claude Opus 4.5 / 4.6 作為基礎模型(根據論文的實驗設定)。因此在 DRACO 上,Perplexity 經過編排的產品(以 Opus 為基礎)會與未經編排的 Anthropic Opus 模型進行 benchmark 比較,並以約 10 個百分點的差距勝出。Perplexity 同時是 Anthropic API 客戶,也是展現其編排層能在 Anthropic 模型之上帶來顯著價值的實體。這是「超越基礎模型的編排」這項發現最明確的實例,也是一項直接反駁模型進步時的 Harness 縮減的案例。
相關連結#
- DRACO Benchmark — Perplexity 撰寫的 benchmark,其產品在上面領先
- 深度研究代理程式 — Perplexity Deep Research 是此系統類別中具代表性的領先實例
- 源自生產環境的評估 — DRACO 的方法:以 Perplexity 自家生產流量打造 benchmark
- Anthropic — Perplexity 以 Claude Opus 4.5/4.6 作為基礎模型,並以未經編排的 Opus 作為 benchmark 對照;同時是客戶與競爭者
- Google DeepMind — 競爭者(Gemini Deep Research 也接受評估);Perplexity 也選用其 Gemini-3-Pro 作為 DRACO 的主要評審模型
- LLM-as-a-Judge — DRACO 的評分方法;Perplexity 透過人類對齊研究選定評審模型
開放問題#
- 供應商發布自家產品勝出的 benchmark,顯然存在誘因問題——DRACO 隨時間推移要如何維持公信力?Perplexity 會實際執行可自動化的更新嗎?
- Perplexity 仰賴 Anthropic(及其他公司)提供基礎模型,同時又在終端產品上與它們競爭——若基礎模型開發者推出自家的深度研究模式,這項編排優勢能維持多久?
資料來源#
- DRACO: a Cross-Domain Benchmark for Deep Research Accuracy, Completeness, and Objectivity — DRACO 論文:Perplexity 為作者;Perplexity Deep Research 為排名第一的系統;Opus 4.5/4.6 為其基礎模型;Gemini-3-Pro 為評審模型
Cited by 7
- DRACO Benchmark×2
Perplexity / Anthropic / Google Deepmind — benchmark author; makers of evaluated systems and the…
- Anthropic
Perplexity — Anthropic API customer and deep-research competitor: runs Opus 4.5/4.6 as base models,…
- Deep Research Agents
Perplexity — builder of the leading deep-research system and of DRACO
- Google DeepMind
Perplexity — deep-research competitor whose DRACO benchmark uses DeepMind's Gemini-3-Pro as…
- Entities — People, Orgs, Tools & Projects
Perplexity — AI answer-engine company; maker of Perplexity Deep Research (the leading system on its…
- Open Questions Backlog
Perplexity ×2 (oldest 106d) — A vendor publishing a benchmark its own product wins is an obvious…
- OpenAI
Perplexity — a deep-research competitor that runs Anthropic (not OpenAI) base models; OpenAI Deep…
Related articles
- DRACO Benchmark
Perplexity's benchmark of 100 production-sourced deep-research tasks (10 domains, 40 countries) graded by 26-expert rub…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
- Deep Research Agents
Agentic systems that decompose a complex query, iteratively search diverse sources, and synthesize a structured, cited…
- Google DeepMind
Google's AI lab; built AlphaProof Nexus; Gemini models, AlphaProof, AlphaEvolve, and the open-weight Gemma line; opens…
- LLM-as-a-Judge
Using one LLM to grade another's outputs against criteria/rubrics; DRACO's protocol is per-criterion binary MET/UNMET +…
