H
Howardism
Plate IIEntities機器翻譯 · machine-translatedENHOWARDISM

Gemini Enterprise Agent Platform

Google Cloud 的代理程式平台:GenAI 評估服務,具備自適應 AutoRaters(與 DeepMind 共同打造)、User Simulator、Automatic Loss Analysis、Online Monitors、OTel 追蹤,以及 ADK/agents-cli 工具鏈;以兩個套件提供 quality-flywheel 評估 skill

Article metadata
Publication details
Published:July 2, 2026
Filed:Entity
Domain:Entities
Tags:EntityPlatformGoogleEvaluation
Reading:4 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Gemini Enterprise Agent Platform 插圖

資料來源#

摘要#

Google Cloud 用於建置、執行及評估代理程式的平台;在此語料庫中,它是 Agent Quality Flywheel 背後的基礎設施。其評估堆疊最值得注意:一項 GenAI 評估服務,其中的 AutoRaters(與 Google DeepMind 共同開發,且依 Google 所述,也用於自家模型與第一方代理程式)是自適應的以模型為基礎的評審——對於多輪代理程式,它們會從對話中擷取使用者意圖、為每個案例產生評分規準、依各項標準驗證追蹤紀錄,並彙整多個樣本進行多數決投票。

提及的元件#

  • GenAI evaluation service — 飛輪最佳化器/評估器分離中的獨立評分器;提供預先定義的多輪 AutoRaters(multi_turn_task_success、multi_turn_trajectory_quality),以及自訂評分規準指標。
  • User Simulator — 在尚無實際流量時,合成多輪情境以進行冷啟動評估。
  • Automatic Loss Analysis — 當失敗判定達十次或以上時,將失敗案例分群。
  • Online Monitors — 持續評估線上正式環境流量,並將品質分數寫入 Cloud Monitoring。
  • OTel tracing — 代理程式會輸出 OpenTelemetry 追蹤紀錄(ADK 預設如此);正式環境追蹤紀錄也可作為評估資料集。
  • ADK (Agent Development Kit) + agents-cli — Google 的代理程式框架與 CLI 工具鏈;adk-samples 代理程式是飛輪的示範對象。
  • The two skill packages — google-agents-cli-eval(ADK/agents-cli)和 agent-platform-eval-flywheel(Evaluation SDK,適用於任何框架),透過 skills.sh 的 npx skills add … 安裝。
  • Model distribution — 此平台是 DeepMind Gemini 3.5 Flash-Lite(2026-07-21)的四個具名發布管道之一,其他還有 Gemini App、AI Studio 和 Gemini API:高效率級模型由此觸及企業代理程式。
  • 交付夥伴承諾(2026-09-08)。 Accenture 與 Google Cloud 表示,雙方新成立的 Accenture Gemini Enterprise Business Group 將建立一支 1,000 人的前線部署工程師團隊,由 Google Cloud 培訓,為 Gemini Enterprise 建置客製化代理式應用;此舉將建立在 Accenture 約 50,000 名具備 Google Cloud 技能的專業人員之上。這是本語料庫中,單一平台部署層所獲得最大的一項人力投入承諾;該公告屬於 vendor-claim,未說明時程、計費模式或雇主(前線部署工程作為交付層、Accenture)。

在語料庫中的定位#

這是 Google 版的代理程式堆疊,對應 Claude Code 與 Codex;但那些條目著重於使用代理程式進行建置,此平台在語料庫中的角色則是衡量代理程式:它將評估功能(評審、模擬器、監控器)打包成產品介面。交付方式是由你現有的任何程式碼代理程式驅動的 skill,本身也證明了 skills-as-distribution-unit 的模式(Agentic Work Systematization)。

相關連結#

資料來源#

§ end
Cited by 6
  • Google DeepMind×3

    Gemini Enterprise Agent Platform — the Cloud product surface where DeepMind-built AutoRaters ship…

  • Accenture

    Gemini Enterprise Agent Platform — the Google Cloud platform the 2026 business group is staffed…

  • Agent Quality Flywheel

    Gemini Enterprise Agent Platform — the platform whose evaluation service, User Simulator, Online…

  • Forward-Deployed Engineering as a Delivery Layer

    Gemini Enterprise Agent Platform — the platform the 1,000-FDE workforce is being staffed against

  • LLM-as-a-Judge

    Google's Gemini Enterprise Agent Platform AutoRaters (developed with Google Deepmind; the grading…

  • Entities — People, Orgs, Tools & Projects

    Gemini Enterprise Agent Platform — Google Cloud's agent platform: the GenAI evaluation service with…

Related articles
  • Agent Quality Flywheel

    Google's eval-fix loop packaged as a skill your coding agent drives: Build & Test → Ship & Monitor → Learn & Refine, ex…

  • Anthropic

    AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…

  • Evals as Product Spec

    Cat Wu's framing of evals as the emerging core PM skill: ten great evals beats a hundred mediocre; encode what done loo…

  • Loop Engineering

    Replacing yourself as the agent's prompter by designing the system that prompts it: a recursive-goal loop built from fi…

  • Open Questions Backlog

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…