H
Howardism
Plate IIProduct & Org中文HOWARDISM

Role Averaging, Not Role Elimination

PublishedJuly 3, 2026FiledConceptDomainProduct & OrgTagsProduct ManagementTeam DesignAI Native OrgRole EvolutionReading18 minSourceAI-synthesised

Andrew Ambrosino's nuanced OpenAI-side take on role collapse: your role is 'the average of what you spend your time on' and tool-gatekeeping is eroding — but eliminating roles dangerously eliminates specialties with knowable best practices ('getting rid of the product role is a terrible idea'), and 'zone defense' coverage plus managers remain necessary because not everyone can work on everything in both breadth and depth

Illustration for Role Averaging, Not Role Elimination

Sources#

Summary#

Andrew Ambrosino's (OpenAI Codex) take on role collapse is the counter-caution to the wiki's mostly-optimistic Engineer PM Convergence cluster. He confirms the convergence — the Codex org has "seen a lot of role collapse," designers "speak engineer," PMs "write code," and "your role is the average of what you spend your time on" (mostly PM work this week ⇒ "you're a PM for now"). But he draws a hard line: eliminating the concept of roles dangerously eliminates the idea that disciplines are specialties with knowable best practices. "I've heard a lot of companies say we're getting rid of the product role — which is, by the way, a terrible idea… that whole discipline of product, with real best practices, real things that have been tried and failed, just gets abandoned because people are like 'oh, I wrote some code.'" The fluidity is real; the flattening-to-mush is a mistake.

Evidence note. practitioner-opinion — one leader's account of his org (OpenAI Codex), explicitly noting it may collapse roles more than other parts of the company/economy because it began as a technical product for engineers.

What erodes vs. what stays#

The distinction Ambrosino draws is between tool-gatekeeping (going away, good) and discipline judgment (staying, essential):

  • Erodes — the "this isn't your lane" boundary and tool-mastery-as-identity. "I spent so long feeling like I shouldn't be a software engineer because I didn't care about assembly language or memorizing TypeScript syntax." Being good at a role got conflated with mastering its tool; that gate is falling. It's now "easier to switch roles, easier to learn the best practices, easier to not tie your effectiveness to the exact tool."
  • Stays — every discipline's real skill component. "Engineering is a skill… other roles aren't just people vibing." His clincher: "You can use Excel, but you cannot work on the finance team." Product, design, and engineering each carry accumulated best practices that surviving-a-vibe-code doesn't confer.

The reconciliation: roles blur at the tool layer and average across a person's week, but the disciplines remain distinct bodies of expertise you can't skip. This is the returns-to-expertise finding restated as org design — expertise stays decisive even as execution cheapens.

Zone defense: coordinating the averaged org#

When "everybody's building everything" and 90 uncoordinated builds appear (Implementation Abundance Inverts Product Work), coordination can't be top-down. Ambrosino's product org runs "zone defense": product people spread out to cover gaps rather than clustering. "If two product people are working too closely, that's often not a good signal." You do a "force-directed activity" — figure out who's best at what, "create space between us so we've got full company coverage," then fill the gaps. The taste-makers "guide things from inception to what the product should be," because "the top-down, year-long planning thing is not going to work" amid the chaos of everyone throwing ideas everywhere.

IC and manager are both, now#

Role-averaging dissolves the IC/manager binary too: "It's not that management is going away or that everyone's an IC — everyone's kind of both." An IC "is not typing code out character by character — you're managing agents, managing work that comes together." A manager of teams "does the same thing at a different granularity." Both are steering; the difference is scope. This is why managers don't disappear: "not everyone can work on everything, in both breadth and depth." (The Anthropic-side Managers as ICs reaches the same both/and from the other direction — managers stay ICs; Ambrosino says ICs became managers-of-agents.)

The member-of-technical-staff thread#

The generic "member of technical staff" title (which Ambrosino traces back to Xerox and research-focused companies, now spreading) is the naming convention for role-averaging: your function isn't fixed by a bucket. But he resists reading it as "functions will disappear" — it's fluidity and easier role-switching, not the end of specialties.

The averaging, surveyed (ICONIQ, Q2 2026)#

Ambrosino's "averaging, not elimination" is one leader's account of his org; ICONIQ's State of AI 2026 survey (~305 AI-building software companies, empirical) puts a population number on the distinction between the two. Asked how AI-driven productivity changed their headcount planning (single-select, N=302): 45% plan "a different mix of roles (e.g., fewer operational, more AI-fluent talent), but not a net reduction in headcount" — the plurality choice, and almost a textbook restatement of this page's title. A separate 33% plan a genuinely smaller team, so elimination is real at the margin, not absent. The composition tilt is visible by function too: R&D/Product/Sales grow while Customer Support and G&A shrink, and G&A operators report removing finance-ops and order-management roles outright — the elimination edge Ambrosino warns is dangerous when it deletes a discipline's best practices, here landing on operational rather than product/engineering specialties. Population-scale, this is exactly his split: roles average and re-mix far more than they disappear, and where they do disappear it is the operational tier, not the disciplines with knowable best practices. (Survey of intent, self-reported forward plans; see AI-Native Organization for the full restructuring cut and the deck's unit economics.)

The Netflix corroboration (Stone, July 2026)#

Elizabeth Stone independently lands on the same both/and from the largest non-lab vantage in the wiki: role fluidity is healthy ("they speak more languages now than they used to"), but "I still see a craft excellence that's really important in the disciplines that I don't think is going away anytime soon… I still find great engineering to be scarce, great data science to be scarce, great creativity to be scarce." Her per-discipline nuclei map onto Ambrosino's specialties-with-best-practices: data scientists own can-we-trust-this-data, PMs own framing the what, engineers own the how. Where she goes further: the specialist end contracts too — fewer narrow deep specialists, more adaptable generalists, with exceptions only for few-people-in-the-world domains (encoding, playback) — see Systems Thinking Over Specialization.

The averaging, randomized: half a role automated, the other half doubled#

ICONIQ surveys intent and Ambrosino describes his own org. Jabarian & Henkel's hiring field experiment is the same split observed under randomization, on one occupation, with administrative outcomes — and it lands on averaging rather than elimination in an unusually literal way.

A recruiter at the partner firm does two expert tasks: conduct interviews and evaluate applicants. The experiment automated the first and left the second untouched, then measured what happened to the second. It got bigger: median interview→offer-decision time rose from 2.62 to 7.24 days, because a recruiter reviewing a conversation they did not have works harder than one who was in it. Headcount was never the outcome — 131 recruiters kept evaluating throughout, and the firm's own framing (endorsed by the authors) is "redirecting recruiter expertise toward evaluation and potentially raising standards in low-entry job assessments."

So the role averaged: the execution half was handed over, the judgment half absorbed the freed time and then some. What is missing from Ambrosino's account, and visible here, is that the residual half is not automatically the pleasant half. The recruiters' own forecasts were pessimistic and wrong in a specific way — 61% expected AI-led interviews to be lower quality, 36% expected lower offer rates, and only 12% expected AI's personal impact on them to be positive (against 47% of applicants). Incumbents holding a discipline's best practices predicted the sign of the effect backwards on the very task they were expert in. That is a real complication for "specialties with knowable best practices survive": here the best practices were interview-conducting, they were the part that automated, and their holders did not see it coming.

What a substitution model does with the same question#

The formal-economics answer to "will your role be eliminated" is structurally the opposite of this page's. Banerjee & Singh's HAT model (arXiv 2607.20781, July 2026, practitioner-opinion — a formal model with no data) puts a linear risk-adjusted cost objective on an allocation simplex, and its Theorem 6 says the optimum sits at an extreme point: the whole task goes to the single lowest-cost agent, human or AI, with mixtures arising only from exact cost ties or from constraints imposed on the problem from outside. Theorem 7 makes the switch abrupt — full substitution or none, no gradual transition. Ambrosino's "your role is the average of what you spend your time on" has no representation in that geometry; neither does the recruiter who lost the interviewing half and kept the evaluating half.

The gap is not a disagreement about the world so much as a modelling choice the paper itself flags. Its limitation 5 concedes the single-task abstraction — real organizations run portfolios of interdependent activities where substituting AI in one task changes human value in adjacent ones — and limitation 4 concedes that extreme-point results are "best interpreted as benchmark structural tendencies rather than literal descriptions of operational firms," since no single agent can meet an organization's aggregate capacity. Averaging is what happens once a role is a bundle of tasks rather than one task: substitution can be extreme-point at the task level and still produce recomposition at the role level. That reconciliation is worth holding, because it means the abrupt-transition prediction and the observed gradual role drift are not actually in conflict — they are claims about different units, and the unit is the whole argument.

Where the two do genuinely conflict is on what survives. This page's thesis is that disciplines with accumulated best practices survive because the practices are real; HAT's Corollary 3 says survival is decided by where two cost curves cross — human compensation convex in skill, AI cost near-flat in capability — so a discipline whose best practices are expensive to acquire is, all else equal, more exposed, not less. The hiring experiment is the one datapoint: the best practices were interview-conducting, they were the half that automated, and their holders forecast the effect's direction backwards.

The averaged role that got narrower#

The averaging thesis has a hidden optimism: the residual judgment half grows. The hiring experiment shows it (evaluation 2.62 → 7.24 days), and Ambrosino's zone defense assumes it. Matsui's solo-authorship study is the same averaging at the level of a single paper — a researcher absorbing the execution roles coauthors used to fill — and it finds the opposite on scope.

Comparing authors' own post-2022 output, solo-writers' content breadth (mean pairwise cosine distance among their own papers) is 0.089 against 0.115 for persistent coauthors — 23% narrower (adjusted −0.028, P < 10⁻¹⁶, n = 35,893), while exploration (distance from their own pre-2022 collaborative centroid) is statistically indistinguishable (+0.001, P = 0.07). The averaged role stays where it was and covers less of it. Which is the point: the coauthor was supplying range, not just labor, and an LLM absorbing the execution layer does not restore it. Whether that reads as scope discipline (do only what you can verify alone) or capacity limit is unmeasured — but "the residual half grows" is not a safe default.

Connections#

  • Community Smells Under AI Adoption — evidence on the specialty side of the equilibrium: discerning AI use is associated with more interaction across knowledge boundaries (β =.252–.259, replicated in three models), so the best-practice erosion warned about here is not visible in the aggregate — with the caveat that the study measures interaction frequency, not retained expertise

  • The Solo-Authorship Rebound — role averaging at the level of one paper: the solo author absorbs the coauthors' execution work and keeps the judgment layer, exactly as this page predicts, while the scope of what they attempt contracts 23% — the outcome the averaging thesis does not anticipate

  • Organizational Complements to AI — where the HAT substitution model lives, with its seven predictions checked against the vault; its extreme-point result is this page's counterpart from formal economics, and reconciles only once a role is treated as a bundle of tasks rather than one task

  • Controlled Variance: AI's Edge as Reduced Dispersion — the averaging measured causally on one occupation: the recruiter's interviewing half was automated and the evaluating half grew (2.62 → 7.24 days of median review per hire), nobody was eliminated, and the incumbents forecast the effect's direction wrong on the task they were expert in

  • Task Crossover — the averaging measured (customer-experience workers spend 77% of occupation-specific AI use on other occupations' tasks) and the caution restated: usage telemetry records what was attempted, never whether it was done well, so it cannot distinguish role expansion from dilution

  • Shared Harness, Differentiated Surfaces — role-averaging cashed out as product architecture: Akshay Nathan's stated reason for merging Codex into ChatGPT Work is that "trying to draw a hard boundary based on who you are is gonna be tough." His own version of the thesis is a T-shape — "AI will enable everyone to become a generalist… but then people will have a specialty" — which lands in the same place as Ambrosino's averaging-with-preserved-craft from a second OpenAI leader, though Nathan frames the specialty as chosen interest rather than accumulated discipline

  • Task Saturation: Broad but Shallow AI Diffusion — the labor-data version: occupations shed a fifth of their tasks to AI collaboration, not their existence, and end-to-end automation is the intent of 6.5% of non-routine-cognitive conversations

  • Systems Thinking Over Specialization — Stone's extension: the discipline nucleus survives, but the narrow-specialist end of the distribution contracts in favor of systems thinkers

  • Elizabeth Stone — the Netflix-scale corroboration of averaging-with-preserved-craft

  • Andrew Ambrosino — articulates the averaging thesis and the specialty-preservation caution

  • Engineer PM Convergence — the Anthropic-side convergence account; this page is its OpenAI-side companion and explicit counter-caution ("don't eliminate the discipline")

  • Managers as ICs — the mirror-image both/and: Anthropic keeps managers coding; Ambrosino turns ICs into managers-of-agents; managers persist in both

  • Planning / Execution Division of Labor — the empirical division ("humans decide what, agents decide how") that role-averaging reorganizes teams around

  • Implementation Abundance Inverts Product Work — the abundance ("everybody's building everything") that makes zone-defense coverage necessary

  • Compute Allocator — the averaged IC as allocator/steerer of agent work rather than character-by-character author

  • Parallel Agent Orchestration — "an IC manages agents" made literal: the fleet the averaged role now runs

  • Returns to Expertise in Agentic Coding — the empirical backbone of "specialties don't disappear": domain expertise still decides success

  • Research Taste as the Human Bottleneck — the rubber-stamping risk applies to zone-defense too: taste-makers "guiding from inception" only works if judgment is real, not approval-by-default

  • AI-Native Organization — the population-scale corroboration: ICONIQ's 45%-plan-a-different-role-mix (not net reduction) is this page's "averaging not elimination" measured across ~305 AI-builders, with the elimination edge landing on operational/G&A rather than the product/engineering disciplines Ambrosino wants preserved

Open Questions#

  • Where is the equilibrium between fluidity and specialty — how much role-averaging before a company loses the accumulated best practices Ambrosino warns about?
  • Zone defense assumes enough high-taste people to cover the whole company; does it degrade in orgs without OpenAI's talent density, collapsing back to top-down planning?
  • Does "your role is the average of what you spend time on" survive performance review and career ladders, or does it fragment them the way Cat Wu flags ("we're sacrificing product consistency")? #oq/source Partially answered: Netflix's move (Systems Thinking Over Specialization) is to leave per-level criteria untouched and add a cross-level AI-fluency overlay, explicitly because the tech shifts too fast to encode per level — one large-org existence proof that ladders survive by absorbing fluidity as an overlay rather than rewriting levels.

Sources#

§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 23
Related articles