H
Howardism
Plate IIEntities機器翻譯 · machine-translatedENHOWARDISM

Perplexity

AI 答案引擎公司;Perplexity Deep Research(在自家 DRACO benchmark 上表現領先的系統)的開發者,也是 DRACO 的發布者;在其編排系統中以 Claude Opus 4.5/4.6 作為基礎模型——同時是 Anthropic 的客戶與 benchmark 競爭者

Article metadata
Publication details
Published:June 15, 2026
Filed:Entity
Domain:Entities
Tags:EntityOrgAI LabDeep Research
Reading:3 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Perplexity 插圖

資料來源#

摘要#

Perplexity 是一家 AI 答案引擎/搜尋公司。在本文語料中,它以 DRACO benchmark 的作者身分出現(arXiv:2602.11685,2026 年 2 月發布,並有一位 Harvard 共同作者),也是 Perplexity Deep Research 的開發者。這套代理式 深度研究 系統在該 benchmark 的每個領域與評分軸向都名列前茅。它是本 wiki 中第一家以自家部署流量打造源自生產環境的評估的供應商。

Perplexity 在本文語料中的功能#

  • Perplexity Deep Research — 深度研究代理程式,會拆解查詢、反覆從多個來源擷取資料,並綜合成附有引文的報告。在 DRACO 上,其標準化分數為 70.5%(以 Opus 4.6 為基礎模型),通過率為 72.8%,領先 Gemini Deep Research、OpenAI Deep Research(o3 / o4-mini),以及未經編排、搭配工具的 Claude Opus 4.5/4.6。其特點是深度研究系統中分數最高、延遲最低,輸入 token 用量也最大(每項任務約 779k)——偏重資料擷取,輸出較精簡。
  • DRACO — Perplexity 從自家 Deep Research 在 2025 年 9 月至 10 月的數千萬筆查詢中抽樣,接著去識別化、擴增、篩選並整理成 100 項由專家依評分規準評分的任務(源自生產環境的評估)。已在 Hugging Face 公開發布。

值得注意的結構性事實:既是客戶,也是競爭者#

Perplexity Deep Research 以 Claude Opus 4.5 / 4.6 作為基礎模型(根據論文的實驗設定)。因此在 DRACO 上,Perplexity 經過編排的產品(以 Opus 為基礎)會與未經編排的 Anthropic Opus 模型進行 benchmark 比較,並以約 10 個百分點的差距勝出。Perplexity 同時是 Anthropic API 客戶,也是展現其編排層能在 Anthropic 模型之上帶來顯著價值的實體。這是「超越基礎模型的編排」這項發現最明確的實例,也是一項直接反駁模型進步時的 Harness 縮減的案例。

相關連結#

  • DRACO Benchmark — Perplexity 撰寫的 benchmark,其產品在上面領先
  • 深度研究代理程式 — Perplexity Deep Research 是此系統類別中具代表性的領先實例
  • 源自生產環境的評估 — DRACO 的方法:以 Perplexity 自家生產流量打造 benchmark
  • Anthropic — Perplexity 以 Claude Opus 4.5/4.6 作為基礎模型,並以未經編排的 Opus 作為 benchmark 對照;同時是客戶與競爭者
  • Google DeepMind — 競爭者(Gemini Deep Research 也接受評估);Perplexity 也選用其 Gemini-3-Pro 作為 DRACO 的主要評審模型
  • LLM-as-a-Judge — DRACO 的評分方法;Perplexity 透過人類對齊研究選定評審模型

開放問題#

  • 供應商發布自家產品勝出的 benchmark,顯然存在誘因問題——DRACO 隨時間推移要如何維持公信力?Perplexity 會實際執行可自動化的更新嗎?
  • Perplexity 仰賴 Anthropic(及其他公司)提供基礎模型,同時又在終端產品上與它們競爭——若基礎模型開發者推出自家的深度研究模式,這項編排優勢能維持多久?

資料來源#

§ end
Cited by 7
  • DRACO Benchmark×2

    Perplexity / Anthropic / Google Deepmind — benchmark author; makers of evaluated systems and the…

  • Anthropic

    Perplexity — Anthropic API customer and deep-research competitor: runs Opus 4.5/4.6 as base models,…

  • Deep Research Agents

    Perplexity — builder of the leading deep-research system and of DRACO

  • Google DeepMind

    Perplexity — deep-research competitor whose DRACO benchmark uses DeepMind's Gemini-3-Pro as…

  • Entities — People, Orgs, Tools & Projects

    Perplexity — AI answer-engine company; maker of Perplexity Deep Research (the leading system on its…

  • Open Questions Backlog

    Perplexity ×2 (oldest 106d) — A vendor publishing a benchmark its own product wins is an obvious…

  • OpenAI

    Perplexity — a deep-research competitor that runs Anthropic (not OpenAI) base models; OpenAI Deep…

Related articles
  • DRACO Benchmark

    Perplexity's benchmark of 100 production-sourced deep-research tasks (10 domains, 40 countries) graded by 26-expert rub…

  • Anthropic

    AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…

  • Deep Research Agents

    Agentic systems that decompose a complex query, iteratively search diverse sources, and synthesize a structured, cited…

  • Google DeepMind

    Google's AI lab; built AlphaProof Nexus; Gemini models, AlphaProof, AlphaEvolve, and the open-weight Gemma line; opens…

  • LLM-as-a-Judge

    Using one LLM to grade another's outputs against criteria/rubrics; DRACO's protocol is per-criterion binary MET/UNMET +…