H
Howardism
Plate IIEntities中文HOWARDISM

Gemini Enterprise Agent Platform

Google Cloud's agent platform: the GenAI evaluation service with adaptive AutoRaters (built with DeepMind), User Simulator, Automatic Loss Analysis, Online Monitors, OTel tracing, and the ADK/agents-cli toolchain; ships the quality-flywheel eval skill in two packages

Article metadata
Publication details
Published:July 2, 2026
Filed:Entity
Domain:Entities
Reading:4 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Gemini Enterprise Agent Platform

Sources#

Summary#

Google Cloud's platform for building, running, and evaluating agents — in this corpus, the infrastructure underneath the Agent Quality Flywheel. Its evaluation stack is the notable part: a GenAI evaluation service whose AutoRaters (developed with Google DeepMind, and per Google the same ones used on its own models and first-party agents) are adaptive model-based judges — for a multi-turn agent they extract user intent from the conversation, generate per-case rubrics, validate the trace against each criterion, and majority-vote across samples.

Components referenced#

  • GenAI evaluation service — the independent grader in the flywheel's optimizer/evaluator split; predefined multi-turn AutoRaters (multi_turn_task_success, multi_turn_trajectory_quality) plus custom rubric metrics.
  • User Simulator — synthesizes multi-turn scenarios for cold-start evaluation before real traffic exists.
  • Automatic Loss Analysis — clusters failure verdicts when failures number ten or more.
  • Online Monitors — continuously evaluate live production traffic and write quality scores to Cloud Monitoring.
  • OTel tracing — agents emit OpenTelemetry traces (ADK does by default); production traces double as eval datasets.
  • ADK (Agent Development Kit) + agents-cli — Google's agent framework and CLI toolchain; the adk-samples agents are the flywheel's demo subjects.
  • The two skill packages — google-agents-cli-eval (ADK/agents-cli) and agent-platform-eval-flywheel (Evaluation SDK, any framework), installed via npx skills add … from skills.sh.
  • Model distribution — the platform is one of the four named launch surfaces for DeepMind's Gemini 3.5 Flash-Lite (2026-07-21), alongside the Gemini App, AI Studio and the Gemini API: the efficiency tier reaches enterprise agents through here.
  • Delivery partner commitment (2026-09-08). Accenture and Google Cloud state that a new Accenture Gemini Enterprise Business Group will establish a 1,000-person forward-deployed-engineer workforce trained by Google Cloud to build bespoke agentic applications on Gemini Enterprise, on top of Accenture's ~50,000 Google Cloud-skilled professionals — the corpus's largest single staffing commitment to one platform's deployment layer, and a vendor-claim announcement with no timeline, billing model or employer stated (Forward-Deployed Engineering as a Delivery Layer, Accenture).

Position in the corpus#

The Google-side counterpart to Claude Code's and Codex's agent stacks — but where those entries anchor building with agents, this platform's corpus role is measuring them: it packages evaluation (judges, simulators, monitors) as the product surface. That the delivery mechanism is a skill driven by whatever coding agent you already use is itself evidence for the skills-as-distribution-unit pattern (Agentic Work Systematization).

Connections#

Sources#

§ end
Cited by 6
  • Google DeepMind×3

    Gemini Enterprise Agent Platform — the Cloud product surface where DeepMind-built AutoRaters ship…

  • Accenture

    Gemini Enterprise Agent Platform — the Google Cloud platform the 2026 business group is staffed…

  • Agent Quality Flywheel

    Gemini Enterprise Agent Platform — the platform whose evaluation service, User Simulator, Online…

  • Forward-Deployed Engineering as a Delivery Layer

    Gemini Enterprise Agent Platform — the platform the 1,000-FDE workforce is being staffed against

  • LLM-as-a-Judge

    Google's Gemini Enterprise Agent Platform AutoRaters (developed with Google Deepmind; the grading…

  • Entities — People, Orgs, Tools & Projects

    Gemini Enterprise Agent Platform — Google Cloud's agent platform: the GenAI evaluation service with…

Related articles
  • Agent Quality Flywheel

    Google's eval-fix loop packaged as a skill your coding agent drives: Build & Test → Ship & Monitor → Learn & Refine, ex…

  • Anthropic

    AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…

  • Evals as Product Spec

    Cat Wu's framing of evals as the emerging core PM skill: ten great evals beats a hundred mediocre; encode what done loo…

  • Loop Engineering

    Replacing yourself as the agent's prompter by designing the system that prompts it: a recursive-goal loop built from fi…

  • Open Questions Backlog

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…