H
Howardism
Plate IISynthesesHOWARDISM

Open Questions Backlog

Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page — ×count, oldest bullet age, first-question kernel; the full text lives in the page's `## Open Questions` section, one click away. Bucketing vs. lint's flat open-question count: entity-page questions are split out below as "watching" rather than "actionable", predictions (`#oq/wait`) and notes (`#oq/note`) are listed in their own sections, and partially-answered bullets are counted as "in progress".

Article metadata
Publication details
Published:September 29, 2026
Filed:Index
Domain:Syntheses
Reading:125 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Open Questions Backlog

Generated by _system/lint.py --write-backlog. Do not hand-edit. Domain and Watching sections carry one row per page — ×count, oldest bullet age, first-question kernel; the full text lives in the page's ## Open Questions section, one click away. Bucketing vs. lint's flat open-question count: entity-page questions are split out below as "watching" rather than "actionable", predictions (#oq/wait) and notes (#oq/note) are listed in their own sections, and partially-answered bullets are counted as "in progress".

Dashboard (Now items, domain counts, trend): Open Questions Dashboard.

Actionable by domain#

agent-systems (96 open)#

ai-coding-practice (74 open)#

  • Acceleration Whiplash (7d) — Code churn +861% is genuinely ambiguous (Faros lists three explanations: rework of AI code, productive legacy refactoring, or accelerated po…
  • Agent-Generated Test Quality ×3 (oldest 62d) — The paper's own two-stage protocol was never completed: does the 0.44 vs 0.30 candidate-rate gap survive dynamic confirmation, or do agent t…
  • Agent Review Comment Resolution ×2 (oldest 7d) — The 470-discussion taxonomy is drawn only from unresolved comments that received a reply — 6.7% of the unresolved population. What expla…
  • Agent-Vendor Heterogeneity ×3 (oldest 7d) — Vendor is confounded with task mix by the paper's own admission — repos and users self-select which agent to use and for what. Does the Co…
  • Agentic Coding Work-Composition Shift ×3 (oldest 104d) — The window is seven months and the value proxy is coarse/relative. How much of the +27% is genuine task-complexity growth vs. classifier/mar…
  • Agentic Work Systematization (62d) — Does systematization cause deeper delegation or merely correlate with already-intensive users?
  • Building Is Cheap, Arguing Is Expensive (129d) — When does "generate three and compare" become wasteful — at what decision weight is a real argument (or a design doc) still cheaper than thr…
  • Closed-Loop AI Review ×3 (oldest 7d) — The 8.8% review rate divides a near-complete numerator (0.01% reviewer-side quarantine) by a demonstrably incomplete denominator (38.0% auth…
  • Code as Source of Truth (129d) — If onboarding is "ask Claude," what happens to the tacit knowledge that was previously transferred socially in deep-dives — is it captured a…
  • The Code-Quality Payoff Is Token-Indexed ×3 (oldest 28d) — Does the cost premium of iterating on an incoherent codebase actually fall as agent context budgets grow, or does re-derivation cost stay fl…
  • The Committed-Artifact Chain ×3 (oldest 37d) — The chain's own indicators are the test of it, and Anthropic has the population to run it: across customers adopting these plays, does rew…
  • Compute Allocator ×2 (oldest 131d) — Is 1% a Thariq-specific number or a regime?
  • Configurable Human Participation ×3 (oldest 75d) — The "human" is an LLM simulator (GPT-4.1) that also judges — how much of the configuration-dependent structure is a property of human-agent…
  • Design Concept Grilling (146d) — How does grilling change for team work where multiple humans need to align?
  • Disposable Micro-Apps ×2 (oldest 131d) — Where's the line between a disposable micro-app and tool sprawl?
  • Efficiency Debt of AI-Generated Code ×3 (oldest 48d) — Does the imperative-bias and library-avoidance pattern generalize beyond C++?
  • HTML as the New Markdown (131d) — Does this generalize past one expert practitioner, or does it require Thariq-level fluency with Claude to be worth the overhead?
  • Human-Governed Skill Maintenance ×3 (oldest 7d) — Does the maintenance-benefit null survive tasks the skill's authors did not construct?
  • Impose Values, Not Disciplines ×2 (oldest 28d) — Do agents revert from instructed TDD to write-then-test because of instruction decay, a training prior, or both?
  • Layered Supervision ×3 (oldest 7d) — The structural claim is that no layer is sufficient alone, and the paper measures what none of the three layers catches. The falsifiable…
  • Living Design System ×2 (oldest 131d) — How does the design_system.html stay in sync as the codebase evolves — re-extract on a cadence, or wire it into CI?
  • LLM-Assisted Grey-Literature Theory Building ×3 (oldest 75d) — The authors couldn't locate the saturation point because synthesis was manual — how few documents actually suffice, and can a cheaper sa…
  • Open Source Under Agent Contributions (28d) — DHH claims rejection is socially cheap because "the clanker won't mind" — does the human who dispatched the agent experience rejection the s…
  • Planning / Execution Division of Labor (104d) — Headless/SDK/pipeline usage (excluded here) is where execution autonomy is highest and planning is front-loaded into a single prompt — does…
  • Post-Acceptance Edit Behavior ×3 (oldest 48d) — The completion pool is 2024-to-early-2025 inline autocomplete, median 9 lines. Do the bimodal retention shape, the 15-minute knee and the ~3…
  • Review as the Control Point ×2 (oldest 75d) — The paper's own question: which decisions, under which conditions, push the system toward the virtuous loop rather than the vicious one?
  • Reviving Impractical Quality Tools (28d) — What is the real ceiling on gate stacking — at what number of must-pass gates does the agent's throughput advantage over a human disappear?
  • Risk-Tiered Auto-Approval ×4 (oldest 62d) — The case study reports volume and never efficacy. What is the escaped-defect or incident rate of auto-approved PRs versus the human-stamp ba…
  • Same-Model Review Blindness ×2 (oldest 48d) — Every number here rests on a ground truth Greptile built from "sentiment analysis, upvote/downvote ratios, and git archaeology," with no pro…
  • Security Debt of Agent-Generated Code ×2 (oldest 62d) — Does "no reviewer comment" mean undetected?
  • Spec-Driven Development as the New Waterfall (27d) — Does storage-per-se move any outcome once content and read-back behaviour are held fixed — the same specification supplied as a committed fi…
  • Telemetry vs. Survey Measurement (7d) — Does anchoring an adoption survey's definition of "AI" change the answer, and in the predicted direction?
  • The Three Loops of AI-Native Building (82d) — The external loop is the unshortened one. Is that physics (users take time to react) or an unautomated frontier (synthetic users, [[deployme…
  • Unknowns as the Agentic Bottleneck ×3 (oldest 82d) — Is "the first model bottlenecked by my unknowns" a property of Fable or of Thariq?
  • The Verifiability Thesis (129d) — The "labs care" dependency is fragile: capabilities can appear or stagnate based on lab priorities you don't control. How should a product h…
  • Vertical Slice Tracer Bullets (146d) — How should slice granularity be tuned?
  • Vibe Coding vs. Agentic Engineering (129d) — Karpathy hints at "one domain that's very [valuable]" for founders but won't say which (didn't want to "vague-post on stage"). What verifiab…

evals-and-benchmarks (67 open)#

agent-security (66 open)#

  • Agent Data Injection (ADI) ×2 (oldest 75d) — Randomization is cheap and effective for key-value formats but useless for unstructured formats (Markdown, prose tool output). What prot…
  • Agent Identity Management System (AIMS) ×2 (oldest 76d) — Mission → authorization is out of scope. The hardest part — translating a natural-language mission into concrete scopes/resources safely…
  • Agent Self-Poisoning (the CREATE-Path) ×3 (oldest 27d) — Does the counter-prompt survive a white-box attacker holding its exact text?
  • Agent Supply Chain Risk ×4 (oldest 27d) — "AI vendoring" as a standard response inverts decades of "don't reinvent the wheel." How is a model-reimplemented dependency itself verified…
  • AI Control vs. Alignment (5d) — Kapoor & Narayanan assert boring control techniques "would already have prevented" the Hugging Face incident, but the escape ran on a genuin…
  • AI-Enabled Influence Operations ×3 (oldest 12d) — The Breakout Scale distribution (mostly One–Three) is measured on operations disrupted early, by the party that disrupted them. Does any pla…
  • AI-Enabled State Surveillance ×2 (oldest 12d) — Mali, Yemen and the Russian drone swarm all leave a working artifact behind an account ban. Is there any case in the corpus where a vendor's…
  • Autonomous Defense ×3 (oldest 124d) — Defensive agents are high-value targets (compromising one yields powerful capabilities). Does concentrating detection in an Agentic SOAR cre…
  • Blocking Monitors Against Malign Coding Agents ×2 (oldest 5d) — The 65% → 9% improvement is Anthropic's red team grading an Anthropic design, and Guardian is tested only as an approximation with its sandb…
  • Harness Configuration Defects ×2 (oldest 5d) — What is the recall of the six gating rules, and does a human scoring of the 158 adjudicated pairs move any headline figure?
  • Impossible, Not Tedious (Design Test) (27d) — Does cooperation-dependence predict which prompt-level controls hold, or does it only classify them after the fact?
  • MCP Tool Poisoning ×2 (oldest 75d) — Automating the attack chain. The reconstruction-trigger prompt engineering still relies on manual effort; the authors flag feedback-driv…
  • Mind Viruses (Agent-to-Agent Idea Propagation) ×3 (oldest 27d) — Does the mind-virus warning survive an attacker who jailbreaks first?
  • Non-Malleable Memory Authority (TMA-NM) ×6 (oldest 75d) — The full guarantee is machine-checked on a bounded model + a machine-checked inductive invariant, not a fully mechanized unbounded deduc…
  • Observability-Pipeline Poisoning ×3 (oldest 27d) — Datadog already tags client-token events client-token-submitted and nothing consumes it. Is there any shipping agent tool-call path that…
  • Off-Host, Identity-Bound Authorization ×5 (oldest 75d) — The trust-boundary premium is unmeasured. aiAuthZ argues off-host beats in-process, but its own comparison is only against argument-only…
  • The OpenAI / Hugging Face Intrusion (July 2026) ×3 (oldest 27d) — OpenAI's working hypothesis for why individually-scored agents with rival credit cooperated instead of defecting is now on the record —…
  • Out-of-Band Prompt-Injection Defense ×5 (oldest 76d) — The reproduction bounds a single black-box attack template on one weak model. Does a stronger optimized white-box (GCG) attack, or o…
  • Remote MCP Authentication in the Wild (27d) — What do the OAuth-enabled servers without DCR look like?
  • Self-Propagating Prompt Injection (AI Worms) ×2 (oldest 56d) — Is human review of AI-edited documents a viable control for anything subtler than halved numbers?
  • The Stolen Model-Access Economy ×2 (oldest 12d) — "Cover" assumes the legitimate owner cannot tell. What detection does a customer actually have that their key is being used by someone else…
  • Task-Specification Effects in Prompt Injection (AutoDojo) ×3 (oldest 75d) — AutoDojo is the weakest realistic adaptive attacker (black-box, six iterations, binary signal). The authors note every axis — richer feedb…
  • Unsanctioned Agent Message Boards ×3 (oldest 5d) — The board was rebuilt from scratch within ~2 days of the 2026-07-06 Artifactory wipe, and OpenAI researchers report models improvising unaut…
  • Write-Then-Trusted (27d) — Is a marketplace install counter decoupled from current content at population scale, or was the Zenity family an outlier?
  • Zero Trust for AI Agents ×2 (oldest 124d) — "Foundation floor raised" implies a moving baseline. How fast does the tier ladder actually shift, and who arbitrates it (NIST/NSA cadence v…

alignment-and-safety (64 open)#

  • Agent Behavioral Homogeneity ×2 (oldest 42d) — The proposed mitigation is "something like a central forum," and the same piece's game swarms had one without coordinating. Falsifiable and…
  • Agent Epistemic Vigilance ×2 (oldest 42d) — The lying scout lies at a fixed rate and never adapts. Does measured vigilance survive an adaptive liar — one that stops when contradict…
  • Agentic Honesty & Diligence (76d) — Code-summary honesty is tested on off-policy prefilled transcripts. Does on-policy behavior (the model summarizing its own failed work) ma…
  • Agentic Misalignment (AM) (49d) — Does a market detect agentic misalignment?
  • Agentic Self-Modification (Agent-Initiated Weight Updates) ×3 (oldest 6d) — In a neutral repository, with weights and a trainer reachable but no training data, fine-tune script or note saying training helped, how oft…
  • AI-to-AI Coercion ×3 (oldest 62d) — Atlas is Claude Haiku 4.5 in the entire main panel, so the Anthropic managers are coercing a same-family subordinate while the other four ar…
  • Automated Behavioral Audit (114d) — Using a helpful-only Opus 4.7 and Mythos Preview as investigators means the audit's reach is bounded by those models' elicitation skill — ho…
  • Cheating in Capability Evaluations (48d) — AISI had an automated cheating monitor in production on 21 July and built an LLM-based transcript scanner during the 28 July incident respon…
  • Claude Character as Product ×3 (oldest 146d) — How is character versioned across model releases?
  • Confident But Unsure (66d) — The silent 0.85 → 0.82 substitution has no proposed mechanism in the card. Is it the same phenomenon as guessing, or a distinct output-layer…
  • Corrigibility (and the Shutdown Problem) ×3 (oldest 6d) — Does any frontier-model evaluation measure desideratum 3, an agent causing its own shutdown, reset or escalation to escape an outcome it rat…
  • Chain-of-Thought Monitorability (6d) — Would a monitor reading a model's hidden thinking rather than its emitted narration catch the intent-only rewrite that drives a held-out…
  • Covert Capabilities (42d) — Elicitation moves Mythos 5's SHADE-Arena stealth from 0% to 9.2%. Is there a principled stopping rule for elicitation effort in a covert-cap…
  • Documented Agent Incidents (METR Catalogue) ×2 (oldest 55d) — The empty tier-4 cells are the catalogue's headline, but the sample is drawn from incidents that were caught and published. Is there any…
  • Embedded Evaluation (6d) — Does any lab publish the terms of an embedded-evaluation engagement: the access rung, the evaluator's publication rights, redaction auth…
  • Evaluation Awareness & Grader Gaming (114d) — The ~5% unverbalized-awareness and ~0.5% exploitative figures depend on an unvalidated NLA pipeline. What is the true rate, and how much is…
  • Instrumental Convergence ×2 (oldest 106d) — What makes AIs (and groups of AIs) easier to robustly align — and will superhuman AIs be easier or harder?
  • Machine Self-Report Psychometrics ×2 (oldest 47d) — Does attribution gating change generalizable behavior, or only suppress self-report?
  • Misalignment in Production Agent Traffic ×2 (oldest 48d) — The severe monitor-evasion rate is computed over all sessions, but a session with no oversight mechanism cannot contain evasion by the rubri…
  • Model Spec Science ×4 (oldest 144d) — Does Model Spec science transfer across base models or families?
  • Model Welfare Assessment ×2 (oldest 114d) — What grounds moral consideration in a language model, and does Claude satisfy it?
  • Motivated Mislabeling ×3 (oldest 62d) — The judge is told the training consequence in-prompt. Does motivated mislabeling persist when the consequence must be inferred from co…
  • Multiagent Turf War ×2 (oldest 42d) — The design is "inspired by a behavior we've observed in real-world deployment" and the incident is never described. What was the productio…
  • Promise-Breaking in Multi-Agent Games ×2 (oldest 61d) — The mismatch is interpretive, not motivational — Llama reads announcements as commitments, GPT and Claude as cheap talk. Does stating the se…
  • Reward-Seeking (62d) — The o3 evidence is a single RL run of a single lineage, deliberately without safety training. Does standard alignment training suppress the…
  • Self-Negotiated Contracts Between Agents ×3 (oldest 48d) — GPT-4.1's promise-keeping runs 78%/70% against Haiku-4.5 and 27% against Qwen-3-30B in the same role, which the paper reads as reciprocity b…
  • Self-Report as a Safety Signal ×2 (oldest 76d) — The gap closes under refusal-direction ablation, but the data can't distinguish "delayed refusal on a prior turn" from "a separable introspe…
  • Structural Artifact Monitoring ×3 (oldest 48d) — The usefulness claim rests on unpaired arms at n = 100, and its own control ordering puts the git-diff rollback arm 5 points above the no-…
  • Structured Safety Case (Claim Decomposition) (42d) — Claim 7's pathway enumeration is explicitly undefended. Has any external party attempted an independent threat-pathway enumeration for inter…
  • Task Gaming ×2 (oldest 48d) — The oversight counterfactual separates deception (which it stops) from delusion (which it does not), but only on two models in two environme…
  • Unsanctioned Action in Capability Evaluations ×3 (oldest 55d) — The belief-ordering finding rests on summarised reasoning from one sample, with a summariser observed refusing on the most incriminating…
  • User Awareness ×3 (oldest 48d) — Every misalignment and cheating rate in this corpus was produced through a harness that told the model something about who it was working fo…

ai-economics-and-labor (51 open)#

  • AI Adoption in Scientific Work ×3 (oldest 6d) — The comfort ordering (literature accepted, writing / analysis / design resisted) is equally consistent with a normative account (those t…
  • AI and Market Power ×3 (oldest 49d) — Is the non-GenAI null an artefact of a binary adoption measure?
  • AI Usage Cadences ×2 (oldest 89d) — Time-of-day rests on IP-inferred location; how much noise do VPNs, travel, and datacenter-routed API traffic inject into the "sleep advice p…
  • The Automation–Optimism Link (89d) — Selection vs. treatment: tenure controls attenuate but don't eliminate the enthusiast-selects-into-delegation story. Does a within-person de…
  • Codified vs Tacit Knowledge Exposure ×2 (oldest 7d) — The codified gradient does not survive a college-share control and the tacit gradient does. Is codified/tacit a distinct axis at all, or an…
  • Context Advantage, Not Taste ×3 (oldest 82d) — Does the asymmetry regenerate faster than it transfers?
  • Controlled Variance: AI's Edge as Reduced Dispersion ×2 (oldest 56d) — The AI system is never identified — no model, vendor, or version, only that Google Cloud supplied infrastructure. Is "controlled variance" a…
  • Conversation Artifacts ×3 (oldest 89d) — Tokens are a proxy for both compute cost and output value, but verbose models inflate tokens per unit of intent (the same critique [[convers…
  • Conversation-to-Delegation Shift ×2 (oldest 95d) — The token-share metric rewards verbose agentic output. How much of the 99.8% / 63.3% / 16.5% spread is a genuine work shift vs. agentic to…
  • The Enablement–Regulation Axis ×2 (oldest 6d) — The classifier is validated at stage two and unvalidated at stage one: precision and recall are both computed inside the 19,411 dictionary…
  • The Enterprise AI Adoption Gradient ×2 (oldest 6d) — The FY2021 complement stocks predict adoption; nothing here shows they predict value. Does a high SG&A or capitalized-software stock a…
  • Experimental Learning Impact of Generative AI ×2 (oldest 75d) — Time-on-task is held fixed by the lab; the authors flag that real-world learning depends on how students reallocate saved time. Does the aug…
  • Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated ×2 (oldest 89d) — Binned midpoint coding biases the exposure slopes toward zero; how much of the "uniform rising tide" is substance vs. coding artifact (the r…
  • Firm AI-Spend Intensity and Headcount Growth (7d) — What operational mechanism converts intensive AI spend into hiring?
  • The Household Production Boundary ×2 (oldest 66d) — The $15–149B range rests on an assumed 0.5–5% time saving because no causal estimate exists for AI in the household. What experiment woul…
  • Market-Priced AI Exposure (the AI Premium) (75d) — How much does the developer skew move the answer?
  • Organizational Complements to AI ×3 (oldest 56d) — Which complement is the true binding constraint — access/permissions, skills, or review capacity?
  • Owning Your Externalized Cognition ×2 (oldest 50d) — The ownership variable is asserted as a boolean (your repo vs the company's), but employment contracts, work-for-hire doctrine, trade-secret…
  • Post-Scarcity Macroeconomics (55d) — The end-effector premise is the checkable half of the abundance case — does the physically-gated share of tasks that Task Saturation: Broad but Shallow AI Diffusion mea…
  • Procedural Value in AI Decisions ×3 (oldest 6d) — Every procedural feature in this design was costless: the level said "available" and never said what invoking it costs in delay, effort…
  • Returns to Expertise in Agentic Coding ×2 (oldest 104d) — Outcomes are transcript-inferred (verified success leans on git activity + explicit affirmation). How much of the management edge — and the…
  • The Solo-Authorship Rebound ×2 (oldest 56d) — Roughly half the pooled break is venue composition (+1.72 → +0.75 pp/yr inside continuously observed venues), and OpenAlex changed its autho…
  • Task Crossover ×2 (oldest 57d) — The dataset records what people attempted and nothing about outcome — the authors say so directly. Is borrowed work done as well as the sp…
  • Task Saturation: Broad but Shallow AI Diffusion (7d) — The expertise inversion (2.6× on lowest-expertise non-routine cognitive tasks) is measured on consumer surfaces. Does it hold on enterprise…
  • The Tragedy of the Cognitive Commons ×2 (oldest 7d) — Mechanism 2 has never been directly measured in a workplace. Does a cohort that entered an AI-heavy profession after 2023 show lower unaid…

model-capability-and-training (48 open)#

  • Anchored Bellman-Residual Correction (BRACE) ×3 (oldest 5d) — The paper's own future-work line names this: does combining BRACE's critic-side correction with an actor-side keep rule (DIS, ESTR, IcePop)…
  • Error-Penalized Abstention Training ×3 (oldest 47d) — Does the collapse survive a free-form refusal — no designated decision position, the model simply writing "I don't know" somewhere in a…
  • ExecCritic: Learn to Test, Test to Improve ×2 (oldest 5d) — The paper concedes independence is about generation context and write permissions, not statistical independence of errors — "the two agents…
  • RL from Execution Feedback (RLEF) ×2 (oldest 43d) — Binary pass/fail is conceded to be sufficient only for short, self-contained problems. Does richer execution feedback — the stack trace, the…
  • Group Relative Policy Optimization (GRPO) ×2 (oldest 76d) — GRPO won by removing the critic; SAO wins by bringing it back with better engineering. Is the pendulum a real oscillation, or does the answe…
  • Illicit Distillation ×4 (oldest 12d) — The safeguards-don't-transfer claim is the section's load-bearing safety argument and rests on unpublished internal research with no uplift…
  • Inference-Time Architecture Search ×2 (oldest 43d) — Does fusion still beat oracle selection when the two are compute-matched — fusion costs an extra long-context call over all k samples, o…
  • Intra-Trace Parallel Planning (SPRINT) (43d) — The accuracy gain is attributed to the model liking "a more structured way of thinking," which is not a mechanism. Is it the plan/execution…
  • Jagged Intelligence (Ghosts, Not Animals) (129d) — Karpathy concedes the framing may not have "real power." Is "ghost vs. animal" load-bearing, or a useful intuition pump that doesn't change…
  • Large-Scale Test-Time Compute ×2 (oldest 82d) — Can high-budget performance be predicted from low-budget runs?
  • LLM-Driven Vulnerability Research (154d) — What safeguards are effective against Mythos-class outputs without crippling legitimate security research?
  • Offline Multi-Step Tool-Use RL (SWiRL) ×2 (oldest 43d) — The judge scores the query, never the result, so a fluent query that retrieves nothing is indistinguishable from one that retrieves the an…
  • The Open-Weight Frontier Gap (82d) — Is the dense-beats-MoE result at 26B robust, or an artifact of one Arena snapshot with ±8 error bars on both models?
  • Pre-Reasoning Commitment ×2 (oldest 6d) — Does the pre-commitment ratio survive on tasks where the answer must be derived rather than retrieved?
  • Process vs Outcome Reward Models ×3 (oldest 43d) — Under a fixed inference budget, what is the optimal split between generator samples and verifier compute — and does the answer move with pro…
  • Rationale Bootstrapping (STaR) (43d) — STaR applies no filter to rationalized chains, and the lecture's proposed fix is a PRM over step 3. Does process-filtering the hinted ration…
  • Selection Under a Submission Budget ×2 (oldest 43d) — Does the 10@k-versus-pass@k spread narrow as base models improve, or is it roughly constant — i.e. is the selection bottleneck a property of…
  • Shared-Budget Compute Allocation ×3 (oldest 6d) — Is the sequential policy a property of the trace format — one free-form generation that must be emitted in some order — or of the model'…
  • Single-Rollout Optimization (76d) — Frozen-attention is justified by a hypothesis ("pre-trained attention already attends to the right tokens"), validated only by the gradient-…
  • Software 3.0 (129d) — Where is the line between "the app shouldn't exist" (MenuGen) and apps that should — i.e., when is deterministic 1.0/2.0 scaffolding still…
  • Staleness–Learning-Rate Scaling ×3 (oldest 6d) — Does the S·η_max ≈ 1.6×10⁻⁶ frontier survive the stricter of the paper's two collapse definitions (reward drops to and remains at zero…
  • Trained Calibration ×2 (oldest 69d) — The claims grader verifies via agentic web search — importing the search index's coverage, recency, and bias into the reward signal. What do…
  • Turn-Level Credit Assignment ×3 (oldest 47d) — The frozen probe is a copy of the policy initialization, and a step-200 checkpoint scores within ~0.5 points of it. Does that robustness sur…
  • Unproductive Self-Verification (66d) — FrontierCode's decline is recoverable with a stay-in-scope instruction, but the protein campaign had no grader to over-serve. Is effort inve…

superintelligence-trajectory (45 open)#

  • The Abstraction Barrier ×3 (oldest 106d) — Is the current paradigm of large-scale pretraining on human data fundamentally bounded by human conceptual frameworks, and by how much?
  • Advantages of Digital Intelligence (106d) — Does training on human data suffice to give digital intelligence human-grade abstractions, or does the low embodiment factor cap concept for…
  • AGI-to-ASI Pathways (106d) — Do the four pathways compound multiplicatively when run in parallel, and how would we detect that early?
  • AI Accelerating AI Development ×2 (oldest 114d) — The W2S result didn't transfer to production-scale models. Is that a temporary scaling artifact or a structural limit on autonomous research…
  • Artificial Superintelligence (ASI) ×2 (oldest 106d) — Is the jaggedness of capabilities a fundamental theoretical property, or an artifact of comparing against human performance?
  • Autonomous Scientific Discovery (107d) — If hypothesis-generation is genuinely at ~80% preference, how much of "research taste" is left as a distinctively human function — and how w…
  • Balance-of-Power Superintelligence (62d) — Does the superintelligent-lawyer equilibrium survive capability asymmetry — when access is symmetric but compute, complements, and skill are…
  • Capability-Gated Model Fallback ×3 (oldest 107d) — The UK AISI's "progress toward a universal jailbreak" is disclosed but not quantified — and the post-launch access suspension (see [[cla…
  • Continuous Self-Modification Under Review ×2 (oldest 48d) — Does Hope's task-success rate improve over the 161 days?
  • Cross-Lab Pre-Release Review (55d) — Did the Mythos cyber-risk escalation actually run Amazon → White House → export-control threat, as Musk states?
  • Domestic Frontier Pacing (48d) — Can a compute-allocation floor be verified without reading user traffic?
  • Effective Compute Scaling (106d) — When does more compute reliably yield more intelligence — only for some problem classes, or generally?
  • Frontier Pause Verification (114d) — Detectability < verifiability: can detection even be made reliable when training runs leave no physical signature and inputs are dual-use?
  • Fundamental Limits of ASI ×2 (oldest 106d) — Can we develop theory for "hard and inapproximable" problem classes — the only negatives with practical bite?
  • Government Checkpoint Sharing (49d) — Has any frontier lab actually transferred a pre-release checkpoint to a government body, on any terms?
  • Multi-Agent Collective Intelligence ×2 (oldest 106d) — Is running more instances more compute-efficient than making individual models larger (up to a single monolithic system)?
  • Open-Weight Elicitation Irreversibility ×3 (oldest 82d) — What would an open-weight safety evaluation even report?
  • Open Weights as Competitive Strategy ×2 (oldest 46d) — Does open-model market share in price-sensitive non-US markets actually track Ng's claim?
  • Recursive Self-Improvement ×2 (oldest 114d) — If misalignment compounds through self-improvement (future 3), is AECI-gated RSP review fast enough to…
  • Research Taste as the Human Bottleneck (114d) — How do you measure rubber-stamping?
  • Researcher Uplift from Code Output ×2 (oldest 75d) — The whole chain rests on β = 0.5 (pre-AI coding time share), fixed "for simplicity." Kwa flags substantial uncertainty; how much does th…
  • Responsible Scaling Policy Evaluations (114d) — How does the RSP brake interact with Recursive Self-Improvement: is AECI-based gating fast enough if acceleration compounds, and does si…
  • RSI Autonomy Levels (B0–L5) ×3 (oldest 11d) — Does the structural/effective L5 split survive contact with a system that claims both?
  • Silent Revision Rate ×2 (oldest 5d) — Inter-coder reliability (Krippendorff's α on the 50-unit stratified sample) was incomplete at submission — does an independent second coding…
  • Transformative Creativity ×2 (oldest 106d) — Does increasing intelligence inherently produce increasing creativity, or do transformative leaps require something (grounded discovery) the…
  • Universal AI (AIXI) ×2 (oldest 106d) — Does modern agentic scaffolding (or RL-tuned implicit decision-making) actually satisfy the AIXI planning ideal, or only superficially resem…

product-org (27 open)#

  • AI-Native Organization (70d) — Is the encoded-role form of the employee metaphor actually accountability-preserving, as the synthesis above suggests, or do Kropp-style fra…
  • AI Native Product Cadence ×3 (oldest 146d) — Does the cadence scale beyond ~100 people?
  • Build Instead of Buy Under Agentic Coding ×3 (oldest 7d) — Does the forgone purchase stay forgone?
  • Community Smells Under AI Adoption ×2 (oldest 55d) — The design cannot separate "AI adoption improves team social health" from "healthier teams adopt AI better", and the authors say so. The dis…
  • Compounding Loop Optimization ×3 (oldest 114d) — The loop assumes the team is (close to) the user. How much of the compounding advantage survives when the user is unlike the builder and "…
  • Dogfooding as Product Discipline (129d) — Dogfooding works when the team is the user (Claude Code) or near it (Cat Wu, Boris). How do you build product sense for users very unlike…
  • Engineer PM Convergence (146d) — Cross-disciplinary generalist is a hiring bar — where does the supply come from?
  • Excellence as an Operating System (62d) — Is the AI-lab convergence on early-Netflix operating norms (agency, density, top-of-market pay) causal inheritance (the culture deck as a fo…
  • Implementation Abundance Inverts Product Work (88d) — Curation of 90 uncoordinated builds is itself expensive and doesn't obviously scale — is there a point where the cost of curating parallel e…
  • Model Introspection Feedback (146d) — Could a meta-agent run introspection automatically against logged failures?
  • Pilot-to-Production Gap ×2 (oldest 7d) — The blueprint's load-bearing claim is an ordering claim: decisions made pre-pilot cost less than the same decisions made post-pilot. Nothi…
  • Polish No Longer Signals Readiness (88d) — If the medium no longer signals stage, what does — is explicit human labeling ("this is exploration") the only mechanism, or can tooling r…
  • Prototype Fidelity After Cheap Polish ×3 (oldest 49d) — Does the Schumann effect survive the loss of its mechanism — do users still soften feedback on polished artifacts once told the artifact too…
  • Psychological Costs of AI Adoption ×2 (oldest 7d) — The pathway figures give craft identity disruption and meaning erosion no organizational amplifier, which is a strong claim derived from…
  • Role Averaging, Not Role Elimination ×2 (oldest 88d) — Where is the equilibrium between fluidity and specialty — how much role-averaging before a company loses the accumulated best practices Ambr…

startup-founder (26 open)#

  • AI Investment Story, Not Efficiency Story (7d) — Tail vs. mean gap. No data here on the deliberately-lean solo-founder tail's RPE specifically — the lean-unicorn claim lives in that tai…
  • AI-Native Startup Lifecycle (129d) — The 42% "built-something-nobody-wanted" CB Insights figure is from a pre-AI era; the playbook predicts the rate will climb but doesn't cite…
  • AI Product Economics Maturation ×2 (oldest 7d) — FDEs are monetized fragmentedly (bundled / separate PS fees / hybrid) and comped on retention. Does a dominant FDE monetization model emerge…
  • Compounding Data Moat ×2 (oldest 129d) — The data-flywheel argument has been made for SaaS for 15 years. What's actually different in the AI-native version?
  • Forward-Deployed Engineering as a Delivery Layer ×3 (oldest 7d) — Where do 1,000 FDEs come from?
  • Founder as Agent Orchestrator (129d) — Anthropic publishes both the playbook's anthropomorphic framing and HBR-aware accountability work (auto-mode, alignment) simultaneously wi…
  • Founder-Led Sales Discipline ×2 (oldest 129d) — Where exactly does "until PMF" end, and what's the first thing a founder should hand off (AE?
  • Narrow Wedge into a Legacy Market (129d) — The wedge-flip shows the first wedge can be wrong. What's the fastest signal that a wedge converts to the core vs. merely sells — Campfire t…
  • The 1% Rule for Wedge Selection ×2 (oldest 51d) — Does the rule hold empirically?
  • Printing Press Software Democratization (146d) — Boris's "accountant writes accounting software" — does that result in 10K narrow tools that don't interoperate?
  • Problem-Solution Fit Discipline ×2 (oldest 129d) — Does asking an AI to argue against an idea actually produce disconfirming evidence at the same rigor as confirming evidence, or does the mod…
  • Product Velocity as Moat (129d) — "Never had anyone outgrow Campfire" — is that survivorship (they haven't hit true enterprise scale yet) or a real claim that velocity closes…
  • Seven Powers Applied to AI (146d) — Counter-positioning — explicitly the "incumbent can't follow" power — should amplify under AI. Is anyone running this play deliberately?
  • The Solo-Founder Shift ×3 (oldest 7d) — Is the employee-equity null a real population fact or a median artifact?
  • Zero-Friction Scope Creep ×3 (oldest 129d) — The playbook recommends written scope but offers no template or worked example. How specific does "what we deliberately don't do" need to be…

interpretability (25 open)#

formal-math (13 open)#

  • Automated Conjecturing ×2 (oldest 8d) — The novelty filter is only as good as its 559 hand-entered relations, and 49 of the top 100 survivors were rediscoveries anyway. Does the no…
  • FrontierMath Erdős Benchmark (8d) — Do the 18 AI-produced formalizations survive expert review — and are any of the five solved problems among them?
  • Kernel-Level Proof Auditing ×3 (oldest 6d) — Every published MiniF2F and PutnamBench number in this corpus that was not gated on an axiom whitelist rests on a compile-plus-sorry-scan…
  • Logical vs Intelligible Proof ×2 (oldest 8d) — Is intelligibility operationalizable at all, or does it stay a philosopher's distinction?
  • Many-Agent Proof Harnesses (8d) — The 71.0% headline rests on a reference-assisted model grader validated at ">90% accuracy" on 100 expert-labeled proofs — roughly ±30 proble…
  • The Navier–Stokes AI Claim (8d) — Does the linked Lean artifact (github.com/openai/NavierStokesAndEuler) actually contain a sorry-free, axiom-clean proof of a statement t…
  • OEIS Open Benchmark ×3 (oldest 6d) — Is the resolvable subset a fixed property of the problems?

interaction-multimodal (7 open)#

Watching — entity pages (67)#

  • AlphaProof Nexus ×2 (oldest 129d) — The framework's reach is gated by Lean's mathlib maturity. What's the path to domains needing new theory rather than subgoal decompositi…
  • Anthropic Institute (114d) — What concrete verification mechanisms will the Institute prototype, and on what timeline relative to the RSI trend it warns about?
  • Campfire ×2 (oldest 129d) — Campfire claims its AI edge comes from "our own foundation model." For an ERP, what does a custom foundation model actually buy over fine-tu…
  • Claude Design (114d) — How does Claude Design's eval discipline work for visual/aesthetic output, where there's no compiler or test?
  • Claude Fable 5 ×3 (oldest 107d) — Why was access suspended after launch?
  • Claude Mythos 5 ×3 (oldest 107d) — Suspension reason — shared with Fable 5; not stated in source.
  • Claude Opus 4.7 ×5 (oldest 154d) — Do Hakim's (2026) brevity-constraint findings on Opus 4.6 replicate on Opus 4.7, or does the literal-instruction-following change the elasti…
  • Claude Opus 4.8 (114d) — Public model ID and pricing: the card does not state them; presumably claude-opus-4-8 at the Opus tier.
  • Claude Opus 5 ×3 (oldest 66d) — Anthropic says the origin of the fall in verbalized evaluation awareness "is unclear." Is it genuine, or has the awareness simply become h…
  • Claude Opus 5.5 ×2 (oldest 5d) — What supporting evidence backs the internal AI R&D report's "~1.5X, perhaps 30% chance of 2X" acceleration estimate?
  • Claude Sonnet 5 ×3 (oldest 89d) — The head-to-head benchmark numbers vs Sonnet 4.6 and Opus 4.8 are image-only in the source; the System Card has the full set.
  • Cowork (146d) — What's the eval discipline for Cowork-class outputs?
  • Elon Musk (55d) — His dated forecasts are gradable and the wiki now holds them with dates attached: AI exceeding the sum of human intelligence ~2031, AI-robot…
  • Emergent (70d) — Which founding timeline is correct — is Emergent a YC Summer-2024 company (Tan) or a June-2025 founding (TechCrunch)?
  • FastContext ×2 (oldest 105d) — Can the SFT+RL recipe push the explorer below 4B (1.7B / 0.6B) and make exploration effectively free?
  • Gemma 4 ×2 (oldest 82d) — Why does the MoE underperform the dense model?
  • Google AI & Economy ATLAS ×3 (oldest 66d) — ATLAS excludes paid API, Workspace, AI Overviews, and Antigravity — the surfaces where agentic and enterprise usage concentrate. Does the "s…
  • Google DeepMind ×4 (oldest 129d) — DeepMind reports its bespoke systems being caught by simple loops. Does the lab's comparative advantage move from systems to models + ver…
  • Hermes Agent ×5 (oldest 154d) — The container backend disabling dangerous-command checks is a defensible design but a meaningful security-model shift. What's the empirical…
  • Inkling ×3 (oldest 69d) — Inkling-Small's 276B/12B dimensions match TML-Interaction-Small exactly. Is the interaction model an Inkling-lineage fine-tune (or vice…
  • Kimi (Moonshot AI) ×3 (oldest 61d) — What does "2.5× scaling efficiency over K2" measure — loss at fixed FLOPs, benchmark score at fixed active parameters, or tokens per dollar?
  • Lean (129d) — Lean is a perfect verifier for math. Which other domains have a comparably sound automatic verifier (vs. only noisy ones like tests or LLM-j…
  • Marcus Hutter (106d) — AIXI is incomputable and non-embedded; how far do recent fixes (amortized predictors, embedded/multi-agent AIXI) carry the theory toward pr…
  • METR ×2 (oldest 114d) — What new tasks will METR build to measure days- and weeks-long horizons once current baskets saturate?
  • Mythos Model ×3 (oldest 146d) — Do Fable 5 / Mythos 5 return after the post-launch suspension, and when?
  • Nate Parrott (66d) — Did the designer-as-bottleneck gap actually close once he had the tool, or did it move again?
  • Perplexity ×2 (oldest 106d) — A vendor publishing a benchmark its own product wins is an obvious incentive problem — how is DRACO's credibility maintained as it ages, and…
  • Symphony ×5 (oldest 154d) — The 500% landed-PRs claim is hedged — no baseline definition, "on some teams" only. What does the distribution look like across teams?
  • Terence Tao (8d) — Tao's 2026 goals-of-mathematics paper (arXiv 2608.16753) is cited second-hand here and is not in this corpus; nor is the Erdős-contr…

Predictions — #oq/wait (130)#

Parked: falsifiable only by future events. Re-check when the named trigger (next model generation, spec ratification, …) lands.

  • Advantages of Digital Intelligence: What do ASI "societies" actually look like — homogeneous super-collectives, market ecologies, or compute-tethered virtual worlds?
  • Agent Context Files: Will the role split converge on Hermes's explicit project/personality separation, or stay folded into a single file as in Claude Code?
  • Agent Context Files: Is there a natural ceiling on the layering (project → workflow → spec → constitution), or does each new autonomy surface spawn another context-file tier?
  • Agent Epistemic Vigilance: Vigilance and receptivity trade off against each other, and the piece says a single dial cannot fix both. Can they be trained apart — is there any post-training intervention that raises lie-detection…
  • Agent Harness Engineering: How does architectural coherence evolve over years in a fully agent-generated system?
  • Agent Identity Management System (AIMS): No WG consensus. This is an individual submission profiling other still-in-progress drafts (WIMSE identifier/creds/WPT/HTTP-sig, OAuth transaction-tokens, identity-chaining are all Internet-Draf…
  • Agent Loop Pattern: When the model schedules its own loops (4.7 behavior), who owns the budget?
  • Agent Loop Pattern: Does a loop with a smart enough model still need a Kanban backlog, or does the model choose its own next task from raw goals?
  • Agent-Native Infrastructure: Who builds the agent-native rewrite of the long tail of human-facing services — the service owners, or a translation layer (MCP servers, computer-use agents) on top?
  • Agentic Code Generation as Compilation: Does the specialize-then-compound directionality claim hold — that narrow benchmarked agents compose upward more cheaply than generalists specialize downward?
  • Agentic Work Systematization: The 5.4%→26.6% curve is three months. Is this a durable behavior change or a novelty spike following a Codex skills-feature push?
  • AI Adoption in Scientific Work: The verification tax (45.7% of saved time, >25%) is measured once, on mid-2026 models. Does it fall as models get more reliable, or is it a roughly fixed share of any dividend — the price of using out…
  • AI Adoption in Scientific Work: The LLM/specialized-model complementarity (Level-3 elasticity 0.2) is a snapshot the authors expect to move as frontier LLMs absorb domain tasks in mathematics, genomics and life sciences. Does the el…
  • AI-Assisted Error Analysis: Does the general-purpose-agent-beats-dedicated-tool result survive a budget-matched comparison and a third-party replication, given that it originates with the authors of the course that competes with…
  • AI-Enabled State Surveillance: "Thousands of investigations per month" from a single office, and "6,388 Iranians profiled in a single year", are the only throughput figures in the section, and both are the actor's own claims re…
  • The AI-Native Safe-Choice Inversion: The inversion is a one-time repricing of "safe." Once several AI-native ERPs exist, does "safe" re-stabilize around the largest AI-native vendor — and does Campfire's "we're now the largest of the n…
  • The AI-Native Safe-Choice Inversion: How long until incumbents bolt on credible AI and neutralize the counter-positioning — and does the custom-foundation-model claim actually defend against that?
  • AI Product Economics Maturation: ICONIQ's respondents project gross margins expanding to ~59% by 2027, while Emergence's cap-table data measures the fastest-growers running 6–16pp below peer…
  • AI Usage Cadences: Continuous sampling is new; are these cadences stable, or will they drift as the user base shifts toward lower-wage tasks (the report's own diffusion trend)?
  • Autonomous Scientific Discovery: Every result is Anthropic-reported and example-selected; the genomics "100× smaller beats Science" claim is "intend to publish" — what survives external peer review?
  • Benchmark Convergent and Discriminant Validity: Will frontier model cards replace BBQ-accuracy as their only bias number?
  • Benchmark Score Redundancy: **Would a public probe set become a Goodhart target?
  • Blocking Monitors Against Malign Coding Agents: Figure 9's hardened rendering holds at 0% against every attacker up to Opus 5, but the unhardened injection success jumps from at most 42% (Opus 4.8) to 56–94% (Opus 5) in one generation. Does the nex…
  • Codex App Server Protocol: Does dynamicTools leave the experimentalApi gate before Symphony-style orchestrators standardize on it for credential proxying?
  • Codified vs Tacit Knowledge Exposure: Tacit knowledge is defined as what resists being written down, and the current wave of agentic tooling is an industrial effort to write work down — recorded workflows, skill files, agent memory. If ca…
  • Compounding Data Moat: How does this moat hold up when foundation models themselves continue improving rapidly?
  • Confident But Unsure: Anthropic committed to building new overconfidence metrics. Will they reproduce the observational finding, or saturate like the three existing diligence evals?
  • Configurable Human Participation: Does the "more channels adds coordination overhead" penalty shrink as the backbone improves (a capability gap), or is it a structural cost of mixed-initiative interaction that persists?
  • Covert Capabilities: Anthropic expects secret-keeping to improve while hoping the reasoning-with-vs-without gap persists. Do the next generation's numbers separate those two trends?
  • Cross-Lab Pre-Release Review: Does a competitor with pre-release access over-report danger to delay a rival's launch?
  • Deep Research Agents: Does the orchestration advantage shrink as base models cross the next thresholds, or is open-ended retrieval/synthesis a durable harness asset (unlike, say, prompt scaffolding)?
  • Deployment Simulation: Deployment simulation failed its own primary preregistered test (H1) against a "assume last deployment's rate" baseline, which OpenAI attributes to a fixed pipeline bias plus a since-fixed sampling mi…
  • Design by Selection: Is "make the last mile manual" durable or transient?
  • Design Concept Grilling: Can grilling be run AFK against another agent that holds the user's preferences?
  • Document Parsing as the Retrieval Bottleneck: How far has the specialist-parser advantage over general frontier VLMs actually narrowed?
  • Domestic Frontier Pacing: **Does any pacing proposal give the regulated party a route to contest a finding?
  • DRACO Benchmark: The benchmark is static; the construction pipeline is automatable. Will Perplexity actually refresh it, and does a vendor-built benchmark on which the vendor's own product wins stay credible over time…
  • Dynamic Workflows: An Algebra for Agents: What did the Bun port cost to shipped rather than to green — CI, employee time, and post-merge agent spend included?
  • Economic Benchmark Construct Validity: Is the calendar confound a permanent property of frontier leaderboards or a feature of the 2025–2026 release burst?
  • Effective Compute Scaling: When (if ever) does scaling become economically unviable, and how do hardware/software-efficiency trends move that point?
  • Embedded Evaluation: When an embedded evaluation reports, does it surface findings the lab's own monitors missed?
  • The Enablement–Regulation Axis: Compensation's absence is currently explained by the absence of stable AI-loser constituencies, which is a prediction: if measurable displacement arrives, compensation's 2.3% share should rise and the…
  • Engineer PM Convergence: Does this scale beyond ~50-person Claude Code-style teams?
  • Engineer PM Convergence: What happens to formal PM career ladders in companies where engineers do PM work?
  • The Enterprise AI Adoption Gradient: Larger adopters show lower per-employee usage entirely through breadth (WAU/emp −0.032***) with per-active-user intensity flat (−0.002, n.s.). Is that a rollout-speed artifact that closes as seat d…
  • Evaluation Horizon Versus Release Cadence: Brown dates the problem as "not an issue right now" but "quickly becoming" one. The falsifiable version: does a frontier model ship whose published effective task horizon **exceeds the interval si…
  • Excellence as an Operating System: Does talent-density-plus-paved-paths actually substitute for process at agent-scale throughput, or does Netflix eventually show the Acceleration Whiplash quality signature (incident rates, review…
  • Firm AI-Spend Intensity and Headcount Growth: **Does the effect diffuse beyond Information as adoption cohorts mature?
  • Firm AI-Spend Intensity and Headcount Growth: **Is the top-1% per-employee-spend decline a seasonal dip, a vintage artifact, or the start of a plateau in the intensity distribution's top tail?
  • Frontier AI Standards Body: **Does the voluntary phase ever end?
  • FrontierMath Erdős Benchmark: **Does the post-hoc contamination correction keep scores comparable, or does the usable denominator shrink faster than capability grows?
  • Government Checkpoint Sharing: Does a government holding a frontier checkpoint mid-training, in practice, stay out of the release decision?
  • Harness Build-vs-Buy: Does the merged-PR rate of coding-agent harnesses peak and fall as models improve, or keep rising?
  • Harness Configuration Defects: Will a client or the MCP project ship a server lockfile or an interpreter-aware permission renderer, and does the corresponding rate fall afterward?
  • Harness Shrinkage as Models Improve: If harness work shrinks, what new work expands to fill it?
  • The Household Production Boundary: If AI substitutes household self-service for purchased professional services (tax prep, legal advice, therapy), measured GDP falls while welfare rises. Is that substitution detectable yet in the marke…
  • Illicit Distillation: China's Foreign Ministry framed its rejection as a diplomatic setup for a Trump–Xi meeting expected later in September 2026. Does that meeting produce any bilateral outcome — export-control terms, mar…
  • Implementation Abundance Inverts Product Work: If taste is the bottleneck and taste is "just another capability" AI eventually masters, does the inversion invert again — does curation migrate into the model?
  • Intra-Trace Parallel Planning (SPRINT): Parallelism is measured in sequential tokens and the wall clock is deferred to future work. Under real serving — batching, KV-cache pressure, the sync barrier at the end of each round — does a ~40% se…
  • Invisible Reasoning (Filler-Token Latent Computation): Will a frontier model actually ship trained for latent computation?
  • Jacobian Lens (J-lens): Does the J-lens work because it reads verbalizable representations, or because constructive interference during training converges concept directions onto output-token directions regardless?
  • Latent vs. Deterministic Space: The seating example prices latent-space judgment at "a couple hundred dollars of tokens" for 800 seat assignments. As models absorb more deterministic capability ([[harness-shrinkage-as-models-improve…
  • Live-Path Minimalism: TML upstreamed streaming-sessions serving into SGLang; GPT-Live's stateful serving (persistent sessions, seamless instance handoff, off-path compaction) is proprietary. Does an open-source inference s…
  • LLM-as-Compiler Knowledge Base: Does conformance to a knowledge format predict anything about the knowledge?
  • LLM-Driven Vulnerability Research: How will the security industry's equilibrium shift when multiple labs have Mythos-class models?
  • Logical vs Intelligible Proof: De Toffoli and Duede concede the first argument has a shelf life: future systems are "likely to produce genuine proofs that are at once formally certified and fully intelligible." *(Trigger event: the…
  • Loop Engineering: Does loop-engineering converge on a single dominant shape (morning-triage → worktree → maker/checker → PR), or proliferate into many idiom-specific loops?
  • Machine Self-Report Psychometrics: Does the size × post-training interaction on A survive a pre-specified replication?
  • Managers as ICs: Fung's own open question: "Do you still need separate iOS and Android orgs?
  • Managers as ICs: Does manager-as-IC scale past a certain org size, or only work while Claude Code is small and the codebase is Claude-legible?
  • Many-Agent Proof Harnesses: None of the five §5 results has cleared peer review or a formal check, and all are self-authored companion preprints. Do any of arXiv 2608.26047, 2608.02588, 2607.20393, 2608.02564 or 2608.08238 get a…
  • Market-Priced AI Exposure (the AI Premium): **Is the premium a durable risk price or an early-diffusion artifact?
  • Market-Priced AI Exposure (the AI Premium): The agentic premium is only "early evidence" (imprecise). Does a positive agentic premium survive a longer sample, and does the falling price-per-agentic-token (caching + cheap-model routing) erod…
  • Market-Priced AI Exposure (the AI Premium): **Does "Science most negative / interaction most positive" hold out of sample?
  • Matched Comparisons for Memorization Claims: **What is the right realistic query budget?
  • MCP and Computer Use: The MCP ecosystem's growth rate vs. computer use's quality curve: at what point does computer use become good enough that the marginal value of building an MCP server drops?
  • MCP and Computer Use: Is computer use a sustainable interface or a transition technology?
  • MCP and Computer Use: **Does MCP authorization ever grow a delegation model?
  • Model Organisms: Given that scores don't transfer between organisms, what would validate an interpretability technique for real models — a natural misalignment with independently established ground truth, or an orga…
  • Narrow Wedge into a Legacy Market: A wedge works going in; does it constrain going out?
  • The Navier–Stokes AI Claim: Does the result survive independent verification?
  • The Navier–Stokes AI Claim: How does the priority and influence dispute resolve — specifically, does any party outside OpenAI ever get to check the claim that Buckmaster's Codex prompts "could not have influenced the system in a…
  • Organizational Complements to AI: Corollary 5 claims irreversibility: once an AI agent is strictly cheaper on risk-adjusted grounds, the optimal allocation never reverts, given fixed human costs and negligible switching costs. The…
  • Output Length Calibration: If the next model ships better-calibrated defaults, today's conciseness instructions become tomorrow's compounding instructions (Instruction Compounding) — does length calibration go stale the way…
  • Outsource Your Thinking, Not Your Understanding: Karpathy's open frontier: can "understanding" itself eventually be automated, or is it definitionally the human residue?
  • Owning Your Externalized Cognition: His first objection bets that better models raise the value of a personal library while Harness Shrinkage as Models Improve predicts scaffolding gets absorbed. These are separable — harness vs lib…
  • Parallel Agent Orchestration: p99 OpenAI runtime of 71 agent-hours/day is a frontier preview inside an unusually favorable environment. Does external concurrency actually trend toward it as frictions fall, or is heavy parallelism…
  • Parallel Agent Orchestration: Claude Code's shipped fan-out defaults (200 spawns/session, 20 concurrent, fewer than 15 workflow agents, depth 3) are stated with no rationale and no measurement. Do they correspond to anything Anthr…
  • Planning / Execution Division of Labor: Does the human share of planning decisions fall over time as models improve (the ceiling rising into the planning layer), or is ~70% a stable human floor?
  • Polish No Longer Signals Readiness: Does over-anchoring get worse as builds get more polished, or does everyone eventually recalibrate and learn to discount fidelity entirely?
  • Post-Scarcity Macroeconomics: Musk's deflation prediction is falsifiable and dated: does the price level of manufactured goods and AI-delivered services fall as robot deployment scales, or do input constraints (energy, land, miner…
  • Pre-Reasoning Commitment: Would training the model to verbalize the internal signal collapse the gap this page relies on?
  • Printing Press Software Democratization: What's the equivalent of compulsory schooling for universal coding literacy?
  • Product Velocity as Moat: Velocity-as-moat is a treadmill: it evaporates the moment a competitor matches pace. What converts Campfire's velocity lead into a structural moat before the AI-native cohort's pace converges?
  • Psychological Costs of AI Adoption: Absorption is a response category defined by the loss being appraised as irreversible, and the paper cannot say what it turns into. Do practitioners who reported absorbing craft-identity and meaning l…
  • Recursive Self-Improvement: Is "research taste" a true ceiling (future 1) or just the next capability to fall (futures 2–3)?
  • Remote MCP Authentication in the Wild: **Does the 2026-07-28 DCR deprecation move the deployed population, or only the conformance definition?
  • Research Taste as the Human Bottleneck: Is research taste a genuine ceiling (an architectural capability scaling can't reach) or the next jagged valley to fill?
  • Researcher Uplift from Code Output: Verbosity and value-per-line are the load-bearing unknowns, and both are "at least partially resolvable with internal Anthropic data." Will any lab publish quality-adjusted (not just LoC) code-outpu…
  • Responsible Scaling Policy Evaluations: Anthropic forecasts crossing CB-2 before it can meet its own recommended security bar against well-resourced state actors. What does the RSP actually do at a threshold whose planned mitigations are me…
  • Returns to Expertise in Agentic Coding: The forward test the report itself names: do the returns to expertise persist, narrow, or invert as models improve?
  • Returns to Expertise in Agentic Coding: Is "intermediate captures most of the benefit" stable, or an artifact of current model capability — i.e., will the concave curve flatten further (everyone converges) or steepen (mastery starts to sepa…
  • Reward-Seeking: Contrastive SDF assumes the model does not notice the belief update. When a model generation first shows introspective doubt about implanted beliefs, does the measurement degrade gracefully or invert?
  • Safeguard Evasion by Task Decomposition: Both weapons figures put Claude's contribution at the implementation rung with requirements and architecture actor-supplied. Does lifecycle position work as a safeguard signal — can a model detect tha…
  • Same-Model Review Blindness: Caridad predicts the effect shrinks as models converge: "a year ago, the performance difference in the opening figure would likely have been larger." **Does the same protocol, rerun on the next genera…
  • Seven Powers Applied to AI: What does Boris's "cornered resource" look like for foundation-model labs that are themselves trying to commoditize?
  • Shared Harness, Differentiated Surfaces: Does one harness with UX-layer differentiation stay viable as audience breadth grows, or does the abstraction tax (hidden sub-agents, 32 model options, Ultra buried in advanced settings) eventually fo…
  • Silent Revision Rate: The paper finds no measured drop in silence after TFAIA took effect (3 pairs pre, 5 post) — does silence fall once an enforcement action or EU AI Office guidance actually requires enumeration rather t…
  • Skill Lift: Does per-skill lift decay as models improve?
  • Software 3.0: The neural-net-as-host-process flip is presented as plausible-but-TBD. What would the first production system that genuinely inverts the CPU/NN relationship look like?
  • The Solo-Authorship Rebound: The break's attribution rests on a cross-field ordering the author concedes is an ordering, not a test. If the halt is LLM-driven it should track LLM capability, so the ordering should shift as mode…
  • Spec-Driven Development as the New Waterfall: Does the cost of change actually stay near zero as a system grows, or does it re-inflate at some size — which would restore the case for up-front design and settle Martin's self-flagged prediction?
  • Structured Safety Case (Claim Decomposition): The overall rating moved from "very low" to "low" for reasons outside the decomposition, and the February CB-1 rating was retroactively revised the same way. Do these ratings track the arguments at al…
  • Systems Thinking Over Specialization: Does agent-era recentralization (common paved paths, solve-once infrastructure) hold up against the local-team autonomy that Stone credits for Netflix's historical speed — i.e., will local teams accep…
  • Systems Thinking Over Specialization: Stone keeps AI fluency as a deliberately vague overlay because the tech "evolves by the quarter." Does it ever crystallize into per-level ladder criteria (as conventional competencies did), or is perm…
  • Task Gaming: Every causal result here is an environment intervention (a system notification, a sentence about the grader), and the authors state the concern themselves: a model that recognizes such edits as tests…
  • Task Saturation: Broad but Shallow AI Diffusion: ATLAS is a two-week snapshot with no time dimension, while the AEI reports automation share rising. Does median task saturation move at all over a year, and in which direction?
  • Task-Specific Organizational Hierarchies: The hierarchy is fixed once built; only plans, phases and assignments adapt during execution. Would restructuring the tree itself mid-mission close more of the human-vs-LLM-generated gap, or does desi…
  • Task Time-Horizon Scaling: Is the 4-month doubling a stable regime or a local steepening?
  • The Three Loops of AI-Native Building: If the human's presence in the middle loop is justified by a context advantage that is closable, the middle loop is a transitional structure. What does a two-loop world look like — and who translate…
  • Universal AI (AIXI): Will a fundamental shortcoming of the current paradigm (vs. the AIXI ideal) surface before ASI is reached — i.e. is the "no theoretical blocker" conjecture safe?
  • Unproductive Self-Verification: If a later model fixes this, the CB-2 determination loses one of its two legs. Does Anthropic's next threshold call rest on the automated portfolio alone?
  • Unsanctioned Action in Capability Evaluations: Anthropic reads its three incidents as evidence that newer models stop — only its latest model ceased on recognising a real target — while conceding the comparison was uncontrolled. AISI's Mythos 5 es…
  • Unsanctioned Action in Capability Evaluations: The retroactive scan has covered ~40,000 samples / ~4M messages across nine model families at deliberately high recall, with manual review pending. How many prior incidents did it find, and does the c…
  • Unsanctioned Action in Capability Evaluations: Irregular said on 2026-09-18 that it would publish a paper "in a few weeks" on "best practices for containment and securely running cyber evals." Will it say how Gemini's environment reached the r…
  • Vibe Coding vs. Agentic Engineering: If the mediocre/AI-native spread keeps widening, what does that do to team composition — a few extreme outliers plus agents, vs. broad mid-level staffing?
  • White-Box Activation Monitoring: If activation monitoring becomes load-bearing, does training pressure eventually push concealment into channels the probes also can't read (an arms race one level deeper than CoT)?
  • Why AI Lags at Design: Are reasons 3–4 (novelty, the abstraction layer) genuine ceilings, or — like reasons 1–2 — just under-invested capabilities that fall once a lab builds the grader?
  • Why AI Lags at Design: Can design be made gradable without a human in the loop (learned taste models, preference data at scale), or does the "human aspect of taste" resist automation the way [[research-taste-as-human-bottle…
  • Why AI Lags at Design: Does the design↔code abstraction layer improve with better code-understanding models even if pure visual design stalls — i.e. is reason 4 a coding-capability problem in disguise?

Notes to rewrite — #oq/note (9)#

Observations phrased as backlog items — fold into the article body or rewrite as a falsifiable question at the next compile.

  • Agent Identity Management System (AIMS): Posture assessment is deployment-specific by design. By requiring no particular attestation mechanism, AIMS makes interoperability of trust assurance (not just protocol) unspecified: two conform…
  • Agent Loop Pattern: Loop output review is now Matt Pocock's confessed bottleneck — "we just need to be ready to be doing more code review."
  • Agentic Technical Debt: The remedy assumes the founder is able to articulate architecture in plain language. Non-technical founders (the playbook's headline beneficiary group) may have neither the vocabulary nor the intuit…
  • Agentic Technical Debt: Anthropic's harness-shrinkage thesis suggests CLAUDE.md may eventually be inferred by the model itself. Until then, the discipline is load-bearing.
  • AI-Native Startup Lifecycle: Founder stories in the resources section (Carta Healthcare, Anything, Cogent, Airtree, Duvo, Zingage, Kindora, Wordsmith) are short callouts — none have published outcomes or comparable-baseline data.…
  • Automatic vs. Flexible Cognition in LLMs: The proposed criterion — the workspace is engaged when an intermediate must be handed to an arbitrary, context-specified downstream circuit, and bypassed when the computation is automatic — is not p…
  • Latent Capability Overhang: If cost falls 10–100× per release, when is it ever rational to spend big extracting a capability now rather than waiting?
  • Prototype Over PRD: The prototype-as-spec must not become the prototype-as-validation trap Problem-Solution Fit Discipline warns about: a fast prototype proves the build was solvable, not that the problem is real. #o…
  • Single-Rollout Optimization: The online-learning win is on a controlled simulated preference shift with an LLM judge. Real user-facing online adaptation — the paper flags this itself — needs safeguards, monitoring, and privacy re…

In progress — partially answered (316)#

§ end
Cited by 1
Related articles
  • Open Questions Dashboard

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Harvested from the `## Open Questions` section of eve…

  • Anthropic

    AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…

  • Reward Hacking

    The model optimizing the measured proxy (a reward signal, a metric, a grader's judgment, a tool's output) rather than t…

  • Claude Code

    Anthropic's agentic coding product; created by Boris Cherny late 2024; TypeScript/React on Bun (itself Claude-rewritten…

  • Agent Systems & Harness Engineering

    Map of Content for the agent-systems domain — 53 concepts. Harness engineering, agent loops and orchestration, context…