H
Howardism
Plate IIAgent Systems中文HOWARDISM

Agent Quality Flywheel

Google's eval-fix loop packaged as a skill your coding agent drives: Build & Test → Ship & Monitor → Learn & Refine, expanded into five stages (prepare data / run inference / grade / analyze failures / optimize); plain-language worry in, metric choice and before/after deltas out; synthetic User Simulator bootstraps, production OTel traces sharpen. Shopify's Sidekick flywheel is the same loop with the terminus moved — instruction and harness edits until they plateau, then production failures mined into SFT+GRPO training signal, so each cycle starts from better weights rather than a longer prompt; its distillation curve crosses the production baseline between 26k and 30k trajectories, and its 96% serving-cost cut is a projection rather than an invoice

Article metadata
Publication details
Published:July 2, 2026
Filed:Concept
Domain:Agent Systems
Reading:25 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Agent Quality Flywheel

Sources#

Summary#

Google Cloud's methodology for engineering agent quality instead of vibe-checking it, shipped (June 2026) as an installable skill that your coding agent drives. The problem it names is the daily reality of agent development: you tweak a prompt, it looks better on three examples, and you have no idea whether you broke ten others — "moved the metric or just moved the vibe." The flywheel is a three-phase loop — Build & Test → Ship & Monitor → Learn & Refine — with the Build & Test phase expanded into five concrete stages, run once in order and then looped (stages 2–5) until quality targets are met. Google states the methodology and its AutoRaters are the same ones it uses on its own models and first-party agents, developed with Google DeepMind. Source is a first-party product blog (vendor-claim): the mechanism descriptions are concrete and the demo cycles are worked in detail, but the results are Google's own demos, not independent measurement.

The five stages#

  1. Prepare Data — build an eval dataset from existing OTel traces, hand-crafted cases, or synthesized scenarios.
  2. Run Inference — execute the agent over the dataset to produce traces (skipped if traces already exist, e.g. production sessions).
  3. Grade — score traces with adaptive AutoRaters (model-based judges that grade a trace and explain why) or custom metrics. The only stage that always runs.
  4. Analyze Failures — read rubric verdicts to understand why a case failed; cluster with Automatic Loss Analysis when failures number ten or more.
  5. Optimize & Iterate — apply a targeted fix, re-run 2–4, compare against the previous baseline.

The skill encodes the discipline that most failing cases take several iterations before metrics actually move — and the architectural rule that the optimizer never grades its own work: whatever proposes a fix (coding agent, automated optimizer, human), an independent evaluation service scores it.

The interface: describe a worry, approve a plan#

The developer never touches the eval CLI and never names a metric. The whole interface is a plain-language concern — "I'm worried about whether travel-concierge honors mid-conversation changes… figure out how to test it and propose a plan" — and the skill's job is to translate that goal into the right evaluation: it reads the agent's code, picks metrics, synthesizes scenarios, runs grading, and reports before/after. This is eval-writing itself being automated: the judgment call moves up a level, from authoring the eval to stating the worry and approving the plan.

In Google's worked demo (an ADK multi-agent trip planner), the skill bootstrapped 25 scenarios with the User Simulator across five revision types, graded with two built-in multi-turn AutoRaters plus a purpose-built categorical rubric, found 21% of revisions IGNORED, located the failure precisely, and — after a three-sentence instruction fix the human approved — re-ran the same evaluation to show 21%→5%.

Promote one concern to a stable metric#

The demo's most transferable lesson. Adaptive AutoRaters regenerate a rubric per case per run, so a specific failure lands as one criterion among several, folded into a blended score — real, named, and invisible: in one case the built-in task-success metric scored a comfortable 0.80 while the user's revision was dropped, because four of five generated criteria passed. Detection is not the problem; isolation is. The move is to promote the one concern to its own stable custom metric — here revision_honored with categorical verdicts (HONORED / IGNORED / PARTIAL / NO_REVISION) — that you can count, gate on ("act if >20% come back IGNORED"), and trend cycle over cycle. The working division of labor: adaptive built-ins as the broad-health signal, one stable measure for the behavior you're changing. (See LLM-as-a-Judge for the adaptive-rubric variant this extends, and Failures That Look Like Success for the failure class the blended score hid.)

It works without a hypothesis too#

Pointed cold at a bug-triage agent with just "find a real failure and fix it," the skill ran broad — varied synthetic scenarios, built-in multi-turn metrics — and surfaced a dominant cluster on its own: in 14 of 15 cases the agent did the work correctly but never told the user which tools it had called (its own instruction requested this; the model had quietly treated it as optional). A one-paragraph fix took tool-disclosure from 0% to 96% in one cycle, per Google's demo. "Here's my goal" and "find me a problem" both land.

Two cadences: dev loop and production loop#

  • On-demand (dev): no real usage yet → the User Simulator synthesizes scenarios. Explicitly a cold-start bootstrap: "synthetic scenarios get you moving; production data is what makes the loop sharp."
  • Continuous (production): the same skill points at real OTel traces — already-complete traces skip Run Inference entirely and are graded in place with the same raters. Online Monitors continuously evaluate live traffic and write quality scores to Cloud Monitoring; when scores drift, failing traces feed the same eval-fix loop. Each production failure is a ready-made test case for the next cycle — Production-Sourced Evaluation operationalized as a product loop rather than a benchmark.

Google's stated direction is to let the skill drive more of the outer loop itself — watching monitors, surfacing regressions, proposing fixes — but today it proposes and a human approves; the shipped version is deliberately not autonomous.

Someone ran that direction anyway. Cline's July 2026 campaign (Agent-Authored Harness Optimization, case-study) executed all five stages unattended for 17 hours — baseline run, failure-trace analysis, targeted fix, re-run, keep-or-discard — with the human doing nothing but pressing continue and reviewing the final PR. Two differences from the flywheel as specified are worth holding onto: the object under optimization was the harness source, not the agent's instructions, and the grader was a fixed benchmark rather than an independent evaluation service, so the optimizer had write access to its own scoring substrate and the invariant was maintained by a prompt clause instead of by architecture.

The same flywheel with the loop closed into weights (Shopify, August 2026)#

Andrew McNamara & Cody Mazza-Anthony (Shopify Engineering, 2026-08-05, case-study) describe the same loop under the same name, running in production on Shopify's GraphQL Sidekick agent — a merchant-facing agent that writes and executes Admin GraphQL API queries against a store, serving up to 2,000 requests per minute. Read it with the COI in front: this is Shopify's own account of Shopify's own system, with no external replication, no adversarial review, and every number self-reported. It sits above Google's vendor-claim post on tier and below every empirical source in the corpus; where it conflicts with a measured result, the measured result wins.

The canonical stage list is the diagram's, not the prose's. The article's "Flywheel loop" figure numbers six stages — 01 Ground Truth Set ("create a ground truth set as specs") → 02 LLM as a Judge ("align LLM Judge to humans") → 03 Frontier Model MVP ("build MVP quickly using frontier models") → 04 Data Collection + Self-healing ("collect production data and mine hard negatives") → 05 Model Training ("fine-tune open-weight models") → 06 Gist Compression ("compress system prompt into Gist tokens") — with a central Flywheel Optimization Process node feeding back to 01. The prose's own ## headings are a different, coarser list that merges 04 and 05 into one section and puts autoresearch where the diagram puts the MVP; both orderings are recorded in the raw file and the diagram's is treated as canonical here.

The structural difference from Google's flywheel is where the loop terminates. Google's Learn & Refine hands a failing trace back to the coding agent, which patches an instruction and re-runs; the artifact that changes is a prompt. Shopify does that first and treats the point where it stops paying as the beginning of the loop rather than the end — verbatim, "once harness improvements plateau, we begin optimizing in parameter space." Their framing of the problem is the sharpest statement of it in the corpus: a deployed frontier model is frozen, so "improvements accumulate in the discrete artifacts around it: prompt edits, retrieval examples, routing rules, and harness code. Production knowledge piles up in words and code while the model's weights remain untouched." Each Shopify cycle begins with a better model; each Google cycle begins with a longer prompt. That is an ordering claim about which object of improvement to exhaust first, made from production rather than from a benchmark — and it is the only such claim in the corpus (see Knowledge-Centric Self-Improvement for the taxonomy of objects it orders, and Agent-Authored Harness Optimization for the harness stage arriving at the same plateau from three other directions).

Stage 01–02: quality is defined before anything else, and the judge is calibrated to a human ceiling#

The article's own emphasis is that this is the step teams rush: "Defining quality is the most important step in the loop… It begins as a specification of what good looks like and becomes the reward signal that drives learning." The concrete protocol:

  • A rubric of a few scored criteria (completeness, execution, response quality, safety) with concrete anchors per score — "your product team's definition of good and bad," and "the quality contract for everything downstream." This is Evals as Product Spec's thesis stated as a production gate.
  • Ground truth must include randomly sampled traffic, not only curated examples: "Golden sets test the cases you already know to look for; random samples reveal what good and bad actually look like in production."
  • A rubric-ambiguity gate before any judge is built: two expert annotators blind-annotate 25 random samples, and inter-annotator agreement is measured with Cohen's kappa. Around κ ≈ 0.2 the rubric is declared ambiguous and rewritten — "if the rubric confuses several product experts who work on this product every day, it will confuse an LLM too."
  • Judge calibration is a prompt-optimization step, not a model choice: DSPy with reflection-based optimizers (GEPA, Agentic Context Engineering).
  • Two validation steps after calibration: backtest the judge against previous A/B tests (can it recover the direction of known wins and losses in engagement or retention?) and run targeted degradation tests — deliberately break one behaviour and confirm the corresponding criterion, and only that criterion, falls.
  • Keep judges small and targeted rather than folding all product behaviour into one, "because focused judges make these tests easier to interpret." Independent arrival at this page's promote-one-concern-to-one-stable-metric discipline.

The load-bearing number lives only in an image. The Judges' agreement figure (headed "even experts do not fully agree") plots human agreement = 83%, annotated "perfect: unreachable," against judge 80% just below it — and neither figure appears anywhere in the article's text. The framing is the one LLM-Judge Validation argues for and the MVVP does not contain: "that agreement is the judge's ceiling… the goal is not a judge that is 'perfect,' but one that matches humans about as well as humans match each other." The honest qualification is that the ceiling comparison itself is reported as raw agreement, not chance-corrected, even though the same team uses κ for the rubric gate one paragraph earlier — see that page for why 83% and 80% are the metric its Finding 1 says overstates.

Stage 03: autoresearch, and the plateau that ends it#

Before touching weights, Shopify pushes the frontier-model MVP as far as it goes "without touching the weights," and does so with an agent rather than by hand. The framing of why prompt tuning is insufficient is precise: the system is "already a production application, with dynamically assembled prompts, custom control loops, and bespoke orchestration spread across a large codebase. No single prompt determines its behavior… The optimization target is the entire harness: its prompts, tool definitions, and orchestration code."

The loop is Karpathy-style autoresearch — propose an edit, evaluate against the calibrated judge, keep it if the score improves, discard it otherwise — configured entirely in one readable program.md whose Environment section grants edit rights to prompts/, tools/ and harness/ with the note "everything here is editable, including the harness," and whose experiment guidance is "small, targeted changes, one idea at a time… prefer the simplest change; revert anything that regresses." Full treatment of this axis, its five other instances and its contradictions on Agent-Authored Harness Optimization.

Stage 04: the self-healing pipeline mines failures into training signal, not test cases#

This is the stage with no analogue in Google's loop. Anonymized production traffic is mined for hard negatives — "conversations the judge correctly scores low and that expose where the model is weakest" — and each one enters a repair pipeline:

  1. A panel of frontier reasoning models critiques the failure independently (the diagram shows three critics on one 2/5-scored conversation: "wrong tool, retry with schema lookup", "missed pagination; fetch all pages", "answer omitted the requested field").
  2. An arbiter merges the critiques into a single repair instruction.
  3. The instruction is injected before the user's turn — "a technique sometimes called 'hinting'" — and the conversation is replayed from that point.
  4. The judge re-scores the replay. Pass → the replay becomes an RL trajectory with the judge's score as its reward. Fail → it is escalated to human annotation (Toloka's expert annotators, scoring against the same rubric that calibrated the judge).

Training then runs in two stages: SFT distillation into a smaller model on complete trajectories including the reasoning that produced them (not just final answers), then GRPO with the calibrated judge as the reward signal. The pipeline runs daily, and the fine-tune is "a full-parameter fine-tune over the accumulated data" — new trajectories plus all previous ones — because "training on both new and previous trajectories limits drift and catastrophic forgetting across cycles."

Two things to hold. The judge is now simultaneously the offline metric, the hard-negative selector, the repair gate and the RL reward — four roles for one calibrated instrument, with a measured 80% raw agreement against an 83% human ceiling, and nothing in the account re-validates it as optimization proceeds against it. Reference-Free Judge Over-Crediting is the corpus's measurement of exactly that failure: the same unchanged judge is valid before optimization and invalid after. And the escalation path means the pipeline's own throughput is bounded by human annotation on precisely the failures the critics cannot repair — the hardest ones.

The distillation curve is the most citable thing in the piece#

The GraphQL Distillation figure is the only place the article reports a scaling relationship, and like the judges' figure its numbers appear nowhere in the text. Axes as printed: x "Dataset Size" (10k–70k), y "Judge Score" (60–76), six labelled points on a rising curve:

Dataset sizeJudge score
13k61.5
26k64.5
30k66.7
39k68.5
46k70.7
61k73.5

Two dashed reference lines, labelled in the chart and not printed as values: "Production baseline" at ≈65.8 and "Frontier reference baseline" at ≈73.1 (both read off the axis). The curve crosses the production baseline between 26k and 30k trajectories and reaches the frontier reference only at the largest size tested, 61k.

Read the top of that curve carefully, because the article's headline rests on it. The chart is titled Distillation — it is the SFT stage alone, before GRPO — and its final point clears the frontier reference by 0.4 judge points, on a chart with no error bars, no seeds, and no confidence intervals anywhere. So the piece's central quality claim ("SFT and RL enable the specialized model to surpass the frontier model performance") decomposes into a measured SFT curve that reaches parity and an un-charted GRPO stage that is asserted to carry it past. The claim is plausible and the curve is genuinely informative about sample efficiency; the "surpasses the frontier" part is not what the figure shows.

Stage 06: gist compression, and what it actually bought#

The last stage attacks serving cost rather than quality. A long static system prompt is "a fixed tax on latency and serving cost, paid on every request," and gisting removes most of it: run the same model as a teacher with the full prompt and a student with a short sequence of learned gist token embeddings, train the embeddings to match the teacher's output distribution with model weights frozen, and ship the student. Shopify reports roughly 6,000 tokens down to about 1,500 with "no measured quality loss on the judge."

Reported effects, all Shopify's own measurements on a 350 requests/minute load test: time-to-first-token down ~19%, end-to-end latency down ~38%, throughput up ~16% more requests per second and ~12% more output tokens per second on identical GPUs, working out to "roughly 14% fewer GPUs for the same traffic." This is a different lever on the same tax Prompt-Cache Economics prices — caching makes a static prefix cheap to re-read, gisting makes it short — and that page carries the comparison.

The economics is a projection, not an invoice, and the article's own modals say so. Verbatim: "Serving this traffic on a frontier model could easily cost an estimated $27M per year based on average token costs. The fine-tuned model could come in at a fraction of the cost, closer to $1M: a 96% reduction in serving cost." Both figures are genuine to the source and both are counterfactual estimates from "average token costs" — Shopify never ran the frontier configuration at 2,000 req/min for a year, and no per-request or per-task denominator is published, so the 96% is a ratio of two modelled annual totals rather than a measured saving (Cost-per-Task Over Cost-per-Token's standing complaint, in a new shape: the percentage is printed and the unit is not). Note also what it does not include: the daily full-parameter fine-tune, the frontier critic panel, the arbiter, the replay passes, and the Toloka annotation contract are all recurring costs of running the flywheel and none appears on either side of the comparison.

The step before stage 1 (Shankar, July 2026)#

The five stages start at prepare data and reach failure analysis at stage 4 — where the skill "reads rubric verdicts to understand why a case failed." That is analysis against a taxonomy that already exists: the rubric named the criteria, and stage 4 explains which one the case tripped. Shreya Shankar's talk is about the step upstream of all five — discovering what the failure modes are, before any rubric can be written — and argues that step is the one an automated pipeline cannot take over, for an epistemic rather than a tooling reason: what "good" means lives in the developer's head, not in the traces.

The two accounts are less opposed than they look, and the seam is this page's own interface. The flywheel's input is a plain-language worry — "I'm worried about whether travel-concierge honors mid-conversation changes" — which is exactly an already-externalized piece of the developer's judgment. Google automates everything downstream of that sentence; Shankar's workflow is about producing it, and her claim is that no tool produces it for you. The cold-start demo ("find a real failure and fix it") is the one case where the flywheel does claim discovery without a stated worry, and it is also where Shankar's objection bites hardest: the tool-disclosure cluster it found was a legible failure, an explicit instruction in the agent's own prompt being ignored. Taste-specific failures are by construction not in the prompt.

Where they agree, and it is the sharper agreement: both put the human at approval rather than execution, and both treat the agent's proposals as suggestions to accept or reject one at a time. Shankar draws the line one notch tighter — her agent applies the human's annotations to other traces but is never asked to extend the taxonomy, because validating an agent's taste costs more than expressing your own.

What it is and isn't#

Is: methodology plus orchestration inside your coding agent — metric selection, eval-service invocation, verdict reading, fix proposals, before/after comparison. Distributed as skills (npx skills add …, two packages against the same evaluation service) — methodology shipping in the same unit as org context, the systematization format crossing vendor lines. Isn't: autonomous (human-in-the-loop); a source of ground truth (AutoRaters are sophisticated but model-based — treat scores as directional, trust deltas between runs more than any absolute number); a substitute for real traffic.

Connections#

  • Optimizer–Evaluator Decoupling — the flywheel's load-bearing architectural rule; proposer and grader stay separate
  • Failures That Look Like Success — the failure class both demo cycles surfaced, and the reason trace-level rubric grading beats output skims
  • LLM-as-a-Judge — AutoRaters are the adaptive-rubric variant of the judge primitive; the flywheel adds the stable-metric-promotion discipline on top
  • Production-Sourced Evaluation — the production cadence: real traces as eval input, synthetic simulation as explicit cold-start bootstrap
  • Evals as Product Spec — the PM skill this automates one level up: the human states the worry and approves the plan; the skill authors the eval
  • Loop Engineering — an eval-fix loop packaged as a product-native skill; Google joining the Codex/Claude Code convergence on shipped loop primitives
  • Agentic Work Systematization — skills as the distribution unit, here carrying vendor methodology rather than org-specific context
  • Compounding Loop Optimization — the same instrument-every-recurring-step discipline, productized for the eval-fix step of agent development
  • Verification as the New Bottleneck — the bottleneck this tooling attacks: grading, failure analysis, and regression comparison made cheap enough to loop
  • Gemini Enterprise Agent Platform — the platform whose evaluation service, User Simulator, Online Monitors, and AutoRaters the skill orchestrates
  • Google DeepMind — co-developer of the AutoRaters
  • Dynamic Workflows: An Algebra for Agents — the flywheel running with no human in the discovery path: Bun's post-merge coverage-guided fuzzers auto-file Claude-authored reproducing-and-fixing PRs (100B parser executions → ~15 PRs), humans review only the output
  • Agent-Authored Harness Optimization — the autonomous version of the outer loop Google says it deliberately has not shipped: same five stages (baseline → run → grade → analyze failures → fix), no human approving each fix, and the object under optimization is the harness rather than the agent's instructions. Shopify's autoresearch stage is a fifth production instance of it, and the only one that reports what happens after the plateau
  • Knowledge-Centric Self-Improvement — the taxonomy of objects an improvement loop can persist into (agent, harness, knowledge, prompt), which Shopify's flywheel extends with a fifth — the weights — and then orders: exhaust the discrete artifacts first, move to parameter space when they plateau
  • LLM-Judge Validation — the validation discipline stages 01–02 partly practise and partly skip: an inter-annotator κ gate on the rubric, an explicit human ceiling rather than 100% as the judge's target, an A/B backtest against the online metric, and per-criterion degradation tests — with the ceiling comparison itself reported as raw agreement rather than chance-corrected
  • Group Relative Policy Optimization (GRPO) — the objective the weight-update stage ends in, here with a calibrated LLM judge as the reward rather than a verifier or a benchmark scorer, running daily on production trajectories
  • Reference-Free Judge Over-Crediting — the risk the self-healing pipeline concentrates: one calibrated judge serving simultaneously as offline metric, hard-negative selector, repair gate and RL reward, with no re-validation as GRPO optimizes against it
  • Prompt-Cache Economics — the other lever on the same static-prefix tax stage 06 attacks: caching makes the prefix cheap to re-read, gist compression makes it short
  • Cost-per-Task Over Cost-per-Token — where the $27M → $1M projection lands: a percentage published without a unit, and a counterfactual annual total on both sides rather than a measured per-request cost
  • LLM-as-Compiler Knowledge Base — the same compile boundary with parameters as the output layer rather than markdown, and the closest production instance of that page's own future direction (fine-tune on your accumulated material so the model knows it in its weights); the distillation curve is what gives that idea a dataset-size number
  • Crystallizing Agent Work into Workflows — the third endpoint for discovered behaviour: promotion into weights rather than into deterministic code. Same evidence-gated keep-or-discard loop on the harness, opposite effect on determinism (a fine-tune buys none, so there is no demotion criterion), and the only source that reports what happens after the harness stage plateaus
  • AI-Assisted Error Analysis — the step upstream of stage 1, and the argument for why it stays human; see the section above

Open Questions#

  • Both demo cycles fixed agents with instruction-level bugs and showed large one-cycle gains. What does the loop look like on failures that need tool, memory, or architecture changes — does "several iterations before metrics move" dominate in practice?
  • The custom rubric is authored by the same coding agent that will later propose fixes. Metric choice is upstream of grading — does decoupling need to extend to who defines the metric, not just who scores it?
  • Synthetic User Simulator scenarios bootstrapped the whole first cycle. How much of the 21%→5% delta survives on real-traffic distributions (the representativeness gap Production-Sourced Evaluation names)?
  • Does a weight-space flywheel actually beat the frontier model it distils from, or only draw level with it? The only curve in the corpus reaches the frontier reference baseline by 0.4 judge points at the largest dataset size tested, with no error bars, no seeds and no intervals, and the claimed overtake rests on an un-charted GRPO stage after it. The falsifiable version: the same distillation sweep with seeds and confidence intervals, reporting the SFT-only and SFT+RL arms separately against the frontier reference.

Sources#

  • Sidekick's continual learning loop — Andrew McNamara & Cody Mazza-Anthony, Sidekick's continual learning loop, Shopify Engineering, 2026-08-05, case-study (~2,100 words, 7 content diagrams). Total first-party COI: Shopify's own account of Shopify's own production system, no external replication, no adversarial review, every figure self-reported. The six-stage flywheel diagram (canonical stage labels), the rubric / Cohen's-kappa / annotator-agreement protocol, judge calibration via DSPy + GEPA/ACE with an A/B backtest and degradation tests, the program.md autoresearch config, the critic-panel → arbiter → hinting → replay self-healing pipeline, Toloka escalation, SFT-on-full-trajectories then GRPO, the daily full-parameter fine-tune over accumulated data, gist compression and the load-test figures, and the $27M/$1M projection. Two figures carry numbers that appear nowhere in the article text and were read directly per the image two-pass rule: Judges' agreement (human 83% / judge 80%, raw agreement) and GraphQL Distillation (the six-point curve and both reference lines, axis labels quoted above). Full parse notes and evidence handling in wiki/sources.md
  • Driving the Agent Quality Flywheel from Your Coding Agent- Google Developers Blog — Melnyk & Dai, Google Developers Blog, 2026-06-30 (vendor-claim); builds on the Cloud Next '26 agent-quality talk
§ end
Cited by 24
Related articles
  • Optimizer–Evaluator Decoupling

    The architectural rule in eval-fix loops that whatever proposes a fix (coding agent, automated optimizer, human) never…

  • Open Questions Backlog

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…

  • Cost-per-Task Over Cost-per-Token

    Anthropic's inverted model-selection default: start with the most capable model and dial effort down — a stronger model…

  • Harness Shrinkage as Models Improve

    Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…

  • LLM-as-a-Judge

    Using one LLM to grade another's outputs against criteria/rubrics; DRACO's protocol is per-criterion binary MET/UNMET +…