H
Howardism
Plate IIEntities中文HOWARDISM

Codex

OpenAI's agentic coding and work platform: a CLI (April 2025) plus a desktop app (built Nov 2025, released Feb 2026) built on the GPT-5-series Codex models, extended by skills/plugins, a headless App Server Protocol, and the Symphony orchestrator; the OpenAI-side reference harness paired against Claude Code, subject of the June 2026 'Shift to Agentic AI' study, and — per its product lead — an app ~90% of OpenAI's whole company uses that is spreading from code into general knowledge work

Article metadata
Publication details
Published:June 26, 2026
Filed:Entity
Domain:Entities
Reading:17 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Codex

Sources#

Summary#

Codex is OpenAI's agentic coding and work platform — the OpenAI-side counterpart to Claude Code throughout this wiki. Released April 2025 as a command-line tool, it grew into a multi-surface agent harness: a threaded interaction model (independent per-task workspaces), reusable skills and installable plugins, a headless App Server Protocol for programmatic sessions, and the Symphony orchestrator that turns Linear into a control plane for it. Originally built for software development — a domain with verifiable, economically valuable, modular outputs — its usage has spread well beyond code into research, drafting, data analysis, and operations.

What it is, in this corpus#

  • The agent harness. Built on the GPT-5-series Codex models; runs multi-step, tool-using, file-modifying tasks. Its threaded model is what enables parallel agent orchestration — many independent agents at once.
  • The systematization layer. Skills (SKILL.md workflow specs) + plugins (installable bundles of skills, MCP integrations, hooks) are the substrate of Agentic Work Systematization; a skill authors a workflow, a plugin distributes it.
  • The headless protocol. The App Server Protocol (JSON-RPC over stdio) drives Codex non-interactively — the basis for orchestration and CI-style use.
  • The orchestrator. Symphony (OpenAI open-source, March 2026) coordinates per-issue Codex workspaces from a Linear board.
  • The usage subject. OpenAI's June 2026 Shift to Agentic AI study measures Codex adoption across individual, organizational, and OpenAI-internal populations — weekly-active usage up >5× in H1 2026, increasingly outside the developer base.

The desktop app (Ambrosino's account)#

The wiki's original Codex entry is the CLI + orchestration stack. Andrew Ambrosino's June 2026 interview describes the desktop app — a distinct surface with its own history and trajectory:

  • Timeline. The team started the app in November 2025, dogfooded it internally, and released it in February 2026. Ambrosino stresses it was a right-sized surface — "sort of a chatbot, but more than that; you could see the code but we weren't going to let you edit it" — deliberately not an IDE.
  • Usage (first-party, unverified). He reports ~90% of OpenAI's entire company (not just engineers) uses Codex, ~100% of employees weekly; 5M+ weekly active users, grown ~6× since January. vendor-claim-tier figures from a product leader.
  • From developer tool to general knowledge work. The pivotal internal discovery: non-engineers (marketing, comms, finance, legal) used the Codex app "even though it is actively hostile to these people" — showing them code, asking to run rg. Attempts to build separate general surfaces failed because "nobody would leave the Codex app." The strategy became one "home base" — start simple, grow complex per user, connect out to specialist tools (it talks to the Excel add-in for finance; opens other apps to finish work) — with the "super app" label Ambrosino says he regrets having to hear about.
  • Self-extension. The signature anecdote: OpenAI's in-house videographer edited launch videos with Codex, which — not being a video editor — built its own Premiere Pro extension to control Premiere by editing the backing files and then talking to the extension it wrote. The agent extends itself into a specialist tool it wasn't designed for.
  • Interaction-modality design. The app juggles connectors, an in-app browser (now on the Atlas "owl" stack with enterprise login), a Chrome-extension bridge, and computer use — Ambrosino calls choosing among them a live, unsettled design problem (keyboard-shortcut mapping, "browser at the top level vs. agent-only browser"). Computer use lets it "just start clicking" through UIs with no API (e.g. the Google Cloud console).
  • Automations as an OpenClaude-style operator. Ambrosino runs scheduled tasks that triage his ~3,000 Slack channels into a daily brief he steers in natural language — an emerging first-class pattern the team wants to make setup-free for non-builders.

The app is also the setting for Ambrosino's product theses: Implementation Abundance Inverts Product Work, "the February app would have failed in November — only the models changed", and Why AI Lags at Design.

The ChatGPT Work merge (Nathan's account, July 2026)#

A month after Ambrosino's interview, Codex stopped being a separate product. Akshay Nathan — who leads product engineering for OpenAI's productivity team — describes the outcome on Latent Space (Codex from 0 to 10M Users: Building ChatGPT Work - Akshay Nathan, OpenAI, practitioner-opinion): ChatGPT Work (launched July 9, 2026) runs on the Codex harness. Full treatment of the architecture at Shared Harness, Differentiated Surfaces; the Codex-specific facts:

  • The harness is literally shared, and only the UX layer is opinionated per surface: git-state visibility (the "dynamic island" assumes a repo), diff-forward chain-of-thought display, and sandboxing defaults. Everything else — plugins, computer use, artifacts, memory, sub-agents, scheduled tasks — is the same in both modes, by principle: "everything that you can do in the Codex portion of the product on desktop, you can do in ChatGPT Work and vice versa."
  • The classic ChatGPT harness still exists for chat (optimized for latency and personality, running the "instant" model); the Codex harness is what knowledge work now runs on, because "if you give the agent access to this infinitely flexible environment as a computer, it can do really powerful things." Nathan frames harness history as "divergence, convergence, divergence, convergence" rather than replacement. Routing between them is a model decision, not a router — "this is the decision that the model is making."
  • Adoption (first-party, unverified). OpenAI said ChatGPT Work + Codex reached 10M combined users within two weeks of the July 9 launch; the counts merged because the harness merged. Nathan calls this "a culmination" but stresses ChatGPT as a whole has hundreds of millions of users, and Work is paid-only, not defaulted on. Earlier in June, OpenAI put knowledge workers at ~20% of Codex's base, growing >3× as fast as developers. All vendor-claim-tier.
  • Codex stays as a brand, deliberately. "Developers have been a core market for us for so long… this doesn't take away from that at all. If anything, it should increase the utility of something like Codex, because now you can move seamlessly between writing a diff to creating an artifact or doing a search."
  • New shared primitives from the Work push: artifacts (agentic Excel/PowerPoint/Docs editing, with a "pretty dramatic" model-side quality jump Nathan attributes to GPT-5.4→5.5→5.6); Sites (hosted HTML, increasingly the canonical team artifact in place of decks and spreadsheets — see HTML as the New Markdown); a persistent computer environment with a file system whose files survive between sessions; and scheduled tasks. Nathan credits OpenClaw for the last two: "it has all the same primitives… scheduled tasks, the ability to store files on a file system, the ability to reference those things over time."
  • Ultra (multi-agent mode) was moved behind advanced settings post-launch and made opt-in, because it "can use more of your limits"; sub-agent transcripts are hidden by default.

The codebase, measured from outside (July 2026)#

The one third-party accounting of Codex-the-repository in this corpus comes from a competitor: OpenHands' twelve-month GitHub analysis (Coding Agents and Technical Debt, case-study). For openai/codex alone, in the year to 2026-07-08: 7,688 merged PRs, 1,202 of them bug fixes (16% — the lowest share of the four harnesses compared), ~3.8M lines changed, ~1.32M lines of current code — the largest merged-PR count and the second-largest codebase in the comparison. The acceleration is the sharper figure: Codex went from ~124 merged PRs/month in mid-2025 to ~1,000/month a year later. And the single-repo layout matches the platform shape described above — "the CLI, dozens of core runtime crates, an app server, MCP support, and Python and TypeScript SDKs" all live in the one repo, so this understates total Codex work wherever it happens privately or elsewhere. See Harness Build-vs-Buy.

The /review feature, traced from outside (July 2026)#

Greptile's research team ran Codex's /review and Claude Code's against 1,000 labelled pull requests (Same-Model Review Blindness, case-study, from a vendor selling a competing reviewer). Two product-level observations, distinct from anything about the underlying model: a Codex review lands at 1–2 comments where Claude Code's lands at 7–8, and the GPT 5.5 traces underneath show the terseness is at least partly prompt-and-post-training, not capability — the model names a deadlock in its reasoning and posts only a lower-severity finding, and recall recovers once an instruction to target 7–10 comments is added. Caridad's diagnosis is that OpenAI's stock /review system prompt aggressively narrows scope to minimize noise, and that Deliberative Alignment makes the model weigh developer-versus-user intent before acting: "the model was not disobedient — it was doing exactly what it was trained to do." Read as a measurement of a shipped default rather than of GPT 5.5.

/goal in production: the Patch the Planet security work (July 2026)#

The corpus's most demanding published use of a Codex feature comes from Trail of Bits, a security consultancy running Patch the Planet — a joint initiative with OpenAI to find and fix bugs in open-source software — with Codex pointed at Rust, curl, zlib and Keycloak (2026-07-28, case-study; the post is co-branded with the campaign and sells both the consultancy's methodology and this product, so weigh the enthusiasm accordingly). Four product-level facts, distinct from anything about the underlying model:

  • /goal is used as a mode, not a slash command. The post says outright that it uses /goal "to refer to goal-based prompting in general," that Codex can set goals for itself through a tool call, and that this is the recommended usage — "we rarely type the slash command ourselves." Several engineers stopped writing goals by hand entirely, handing Codex the threat model and asking it to draft the goal prompt.
  • Codex builds its own security infrastructure. Trail of Bits' summary claim: it "can create custom security infrastructure that takes a security researcher weeks to build in under a day." An orchestrator that downloads every rust-lang/rust P-critical issue as JSON and spawns one Codex session per issue is their worked example.
  • A tool exists because Codex skips reading. They built and open-sourced aicov (github.com/trailofbits/aicov), which tracks what lines of code Codex has actually read, because "Codex has a tendency to skip reading the entire codebase even when explicitly asked," so it "can't 'cheat'." A shipped instrument for a named behavioral defect in this harness.
  • One outcome per session is the operative constraint. Competing outcomes in a single goal ("find bugs" and "achieve high coverage") optimize unevenly; their fix is a separate Codex session per attack surface plus one open-ended roamer. This is the practical form of the threaded model listed above.

The findings this produced — a rustc soundness hole and a miscompilation patched in Rust 1.98, two potential Keycloak SAML privilege escalations, 11 Semgrep CVE-variant hits — are treated at LLM-Driven Vulnerability Research, and the goal-design rules at Loop Engineering. Note what the post does not report about the product: no run counts, no token or cost figures, no false-positive rate for the two-judge validation pass, and no /goal failure modes beyond prompt design.

Codex vs. Claude Code#

The two are the wiki's reference harnesses, repeatedly compared. Loop Engineering's central structural claim is that both now ship the same five primitives (automations, worktrees, skills, connectors/plugins, sub-agents) under different names, so the same agent loop works in either — evidence for harness shrinkage. Where they diverge is institutional: Codex sits inside OpenAI's GPT-5 ecosystem and the Symphony/App-Server orchestration stack; Claude Code inside Anthropic's. The "harness engineering" framing (OpenAI, April 2026) is Codex's house philosophy for an agent-first workflow.

The reviewer role, from outside both vendors (DHH, August 2026)#

DHH (David Heinemeier Hansson) uses the two harnesses in fixed roles rather than choosing between them: Claude Code drives, Codex at xHigh reviews, every time. "I'll have a Opus or Fable do the work, and then I always end it, review with Codex xHigh… and it keeps finding stuff" (Lex Fridman #501, 2026-08-26, practitioner-opinion). He also runs OpenCode as the harness for open-weight models against Fireworks inference, so the three harnesses are differentiated by job — driver, reviewer, open-model host — not ranked. See Same-Model Review Blindness for the measured reason a cross-family reviewer might beat a same-family one.

Connections#

  • OpenAI — maker; Codex is OpenAI's agent-tooling thread in this corpus
  • Claude Code — the Anthropic-side peer harness Codex is compared against (same five loop primitives, different ecosystem)
  • Symphony — OpenAI's open-source orchestrator that drives Codex from Linear
  • Codex App Server Protocol — Codex's headless JSON-RPC protocol
  • Conversation-to-Delegation Shift — the June 2026 usage study built on Codex telemetry; its adoption curve and token-share data
  • Agentic Work Systematization — Codex's skills/plugins are the systematization substrate that study measures
  • Parallel Agent Orchestration — Codex's threaded model is what enables the concurrency the study documents
  • Loop Engineering — Codex as one of the two tool surfaces that now ship all five loop primitives
  • Harness Shrinkage as Models Improve — Codex absorbing harness capability (skills, automations, worktrees) into named product primitives
  • Shared Harness, Differentiated Surfaces — the architecture of the ChatGPT Work merge: one harness, three UX differentiators, and why OpenAI rejected the separate-products shape Anthropic chose
  • Andrew Ambrosino — product & engineering lead for the Codex desktop app; the source for the app's history, usage, and general-knowledge-work pivot
  • Implementation Abundance Inverts Product Work — the product-process thesis Ambrosino draws from building Codex
  • Build for the Next Model — the Codex app as the case study: same shape, different-intelligence releases (Nov→Feb; Operator→Atlas→Codex)
  • Why AI Lags at Design — Ambrosino's design-capability read, developed while building the app's front end
  • Role Averaging, Not Role Elimination — the "role collapse" the Codex org saw more of than the rest of OpenAI
  • Harness Build-vs-Buy — the outside measurement of the Codex codebase (7,688 merged PRs/year, ~1.32M lines, ~124 → ~1,000 PRs/month) and what it implies about forking any harness
  • OpenHands — the open-source competitor that published the comparison
  • LLM-Driven Vulnerability Research — the harness under the corpus's only production vulnerability-research pipeline: one /goal session per Rust P-critical issue, two different-model judges, and a soundness hole plus a miscompilation patched in Rust 1.98
  • Loop Engineering — /goal as this page's contribution to the loop primitives, and the three production rules for writing one (let the model draft and red-team the goal; define the outcome, never the path; one outcome per agent)
  • Same-Model Review Blindness — Codex's /review as one of the two measured reviewers, and the finding that constrains it: GPT 5.5 is at its weakest on Codex-authored PRs (50.5% high-severity recall against 62.0% on Claude-authored ones), so the harness's own review command is least sensitive to the code the harness itself writes

Sources#

§ end
Cited by 44
Related articles
  • Claude Code

    Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…

  • Harness Shrinkage as Models Improve

    Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…

  • OpenAI

    AI lab and maker of the GPT-5 series and Codex; in this corpus it appears as a frontier-safety research source (Deploym…

  • Verification as the New Bottleneck

    Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…

  • Parallel Agent Orchestration

    One human overseeing a team of concurrent agents: OpenAI Codex telemetry's first hard numbers (28.6% of staff peaked at…