H
Howardism
Plate IIEvals & Benchmarks中文HOWARDISM

Cognitive Capability Profiling for Task Suitability

Prunty et al. (Cambridge CFI) put AI systems and workplace tasks in one cognitive space: rubric-annotate 19,535 benchmark items for their demand on 18 capabilities (16 survive an inter-rater screen, clustered to 8), infer latent capability by Bayesian IRT where difficulty is *annotated* not estimated, and ask 410 workers to spend 100 points across capabilities per activity. Six systems share one profile shape — Semantic Memory 5.59, Social Cognition 4.08, Language 4.02 on top; Action Planning 1.99, Instrumental Reasoning 1.22, Object Permanence 0.29 at the bottom — and what separates the leaders is planning and control, not knowledge. A 5.30-wide dimension spread against a 1.12-wide system spread is the whole 'dimensions over families' claim; no variance decomposition is reported. Measures task *importance*, not demand, and is validated against no deployment outcome

Article metadata
Publication details
Published:September 23, 2026
Filed:Concept
Domain:Evals & Benchmarks
Reading:35 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Cognitive Capability Profiling for Task Suitability

Sources#

Summary#

Every AI-exposure instrument in this wiki describes work in the vocabulary of work — O*NET tasks, detailed work activities, occupation titles — and then asks whether a model can do it. Prunty, Tešić, Quinn, Hernández-Orallo & Cheke (Leverhulme Centre for the Future of Intelligence, Cambridge; UK DSIT; Universitat Politècnica de València — [[raw/cognitive-capability-profiles-ai-task-suitability|arXiv 2608.25623]], 2026-08-26, 44pp, empirical) do something structurally different: they describe both the model and the work in the vocabulary of cognition, on one shared set of dimensions, so that the two can be compared without either being described in terms of the other.

The design goal is decoupling. Capability profiling and requirements gathering are separate measurement processes over a common capability space, so each can be re-run on its own cadence — models change fast, what a job cognitively demands changes slowly — and the framework survives a model generation without re-eliciting anything from workers. The authors are explicit that "the principal contribution of this work is therefore methodological."

Evidence note. empirical — 19,535 annotated benchmark items, six systems fully evaluated, 410 post-QC questionnaire respondents, MCMC convergence reported. The tier is kept, but read the What it does not measure section before treating any suitability score as a prediction: the requirements half elicits importance, not demand, and nothing here is validated against a realized deployment outcome.

The three stages#

1 — Cognitive capability profiling. Infer a system's latent capability levels from its pass/fail record on a benchmark battery whose every item has been annotated for the cognitive demand it places on each dimension.

2 — Task requirements weighting. Ask domain experts (here, employees) how important each of the same capabilities is for each of their work activities.

3 — Suitability mapping. Combine: suitability is an importance-weighted aggregation of the agent's capability levels.

The two streams are deliberately different in kind — profiling yields capability levels, requirements yield importance weights — and that asymmetry is the framework's central limitation, taken up below.

Stage 1: annotate the demand, don't estimate it#

Capabilities and rubrics#

18 core cognitive capabilities drawn from the psychometric and cognitive-science literature (Carroll's factor-analytic survey; Spelke & Kinzler's core knowledge), in four families plus three domain-general capacities (Table A1, verified):

FamilyCapabilities
Memory systemsEpisodic Memory, Semantic Memory, Procedural Memory, Prospective Memory
Executive controlWorking Memory, Attention & Inhibitory Control, Cognitive Flexibility, Planning
Object & space understandingPerception & Pattern Recognition, Functional Perception, Spatial Reasoning & Navigation, Object Permanence
Social & communicativeTheory of Mind, Emotion Perception & Empathy, Language
Domain-generalMental Simulation, Metacognition, Causal Reasoning

Each capability gets a six-point demand rubric, 0 (not required) to 5, with concrete anchoring descriptors and example tasks at every level. The levels were motivated by a population intuition — level 1 a demand most adults meet, level 5 one only a small minority do — but the inference depends only on their ordinal structure and geometric spacing, and the paper makes no claim that they correspond to fixed population proportions. Demand levels are defined within each capability: a level-4 Working Memory demand and a level-4 Theory of Mind demand are not assumed equivalent in absolute terms.

This is the methodological pivot the rest of the pipeline hangs on. The rubric method is inherited from the ADeLe framework (Zhou et al., General scales unlock AI evaluation with explanatory and predictive power, Nature 652, 2026), which "unlocks" existing benchmarks for construct-based evaluation by having an LLM judge apply expert-written rubrics to individual benchmark items. It buys the scale of benchmarks and the construct validity of psychometric design at once — and it means per-item difficulty arrives from a rubric rather than from a response population.

The battery#

22 benchmarks, 19,576 items (Table 1, verified cell-for-cell; the item counts sum exactly to the printed total). Target size ~20,000, constructed by blending the observed demand distribution (weight 0.7) with a flat distribution (weight 0.3) — the former keeping coverage representative of common demands, the latter guaranteeing rare high-demand items survive — under a constraint that no single dataset supply more than 10% of the battery. The heaviest contributors are Fantom (2,177), OpenToM (2,000), PlanBench (2,000), AGIEval (1,896), StepGame (1,563), Text Navigation (1,413), MetaMedQA (1,373), BigBenchHard (1,265), EmoBench (1,200); MMLU-Pro contributes 200.

Two annotators, and what the screen removed#

Every item was annotated independently by GPT-4o and Gemini 3 Flash against the rubrics. Inter-rater reliability (Table A2, verified) killed two dimensions outright:

  • Attention & Inhibitory Control: ρ = 0.09, κ_w = 0.07
  • Prospective Memory: ρ = −0.03, κ_w = 0.15

The 16 retained dimensions run ρ = 0.38 (Mental Simulation) to 0.81 (Theory of Mind), cleanly separated from the excluded pair, and even the weakest agree to within one demand level on ≥ 74% of items. Ratings were averaged; 41 items with an invalid rating from either rater were dropped, leaving 19,535.

That two of eighteen capabilities cannot be annotated reliably by two frontier models reading the same rubric is itself a finding, and the two casualties are both executive/temporal constructs — sustained attention and remembering-to-do-later — which are precisely the capacities an item's static text gives an annotator least purchase on.

Eight dimensions, because the demands collide#

Cognitively demanding items recruit many abilities at once, so the demand columns are strongly positively correlated and the inference model has little basis for separating the corresponding latent abilities. Hierarchical agglomerative clustering on the z-scored demand-profile correlation matrix, cut at distance 1 − r = 0.5, yields eight composite dimensions (Table A3):

DimensionConstituents
Episodic Memory (EM)—
Semantic Memory (SM)—
Object Permanence (OP)—
Language (L)—
Social Cognition (SC)Emotion Perception & Empathy, Theory of Mind
Instrumental Reasoning (IR)Causal Reasoning, Functional Perception
Information Integration & Control (IIC)Working Memory, Cognitive Flexibility, Metacognition, Perception & Pattern Recognition
Action Planning & Simulation (APS)Planning, Procedural Memory, Mental Simulation, Spatial Reasoning & Navigation

The paper is careful that the clusters reflect co-occurrence of demands across items, not conceptual similarity — conceptually distinct capabilities may simply always be required together. That is the same identifiability problem Benchmark Score Redundancy hits at the score level, arriving here at the item-annotation level instead, and the resolution is the same one: merge the collinear columns rather than pretend they are separable.

Two properties of the resulting demand distribution matter downstream (Table 2, verified): no dimension reaches level 5 on any item, and coverage varies enormously — IIC and L on 100% of items, SM 99%, APS 94%, IR 92%, but EM 50%, SC 42%, OP 36%. Language is required by every item yet 97% of those are at levels 1–2. Full coverage with a narrow range is an identifiability hazard, which is what the recovery analysis was built to test.

Inference: IRT with the difficulty given, not fitted#

Profiles are estimated with Measurement Layouts (Burden, Voudouris, Burnell, Rutar, Cheke & Hernández-Orallo, 2023), a Bayesian item-response model. Each item carries a demand vector D_j ∈ {0..5}^K; each agent a log-capability vector c_k = log θ_k. Difficulty is δ_jk = e^(λ D_jk) with λ fixed at 1, and the agent's margin on dimension k is the log-ratio m_jk = c_k − λD_jk. Inactive dimensions (D_jk = 0) contribute exactly zero and drop out. Margins pool into an item logit through a sigmoid to a Bernoulli outcome.

The pooling rule is the substantive modelling choice, and it is a single parameter on a continuum from compensatory (strength on one ability offsets a shortfall on another) to bottlenecked (the weakest required capability limits performance regardless of strengths elsewhere). A soft-minimum with temperature τ recovers mean pooling as τ → 0 and weakest-link as τ → ∞.

This is the sharpest contrast with Item Response Theory for LLM Benchmarks. ATLAS calibrates 3PL item difficulty b from the response patterns of ~4,000 Open LLM Leaderboard models, which makes b a property of that examinee population — "hard for the 2023–2024 open-weight fine-tune ecosystem" — and leaves the scale vulnerable to temporal drift (ability MAE degrading ~50% across a single one-year calibration boundary). Here difficulty never touches a response matrix: it is read off a rubric by an annotator that is not the subject. The item bank cannot go stale relative to a model population, because it was never anchored to one. The price is paid elsewhere, and the paper is unusually candid about where.

What the recovery analysis actually establishes#

Twenty synthetic agents sampled from the prior, responses simulated on the real battery, profiles re-fit. Two choices fall out (Tables 4 and 5, both verified cell-for-cell):

Pooling temperature. Under the compensatory baseline (τ = 0) the three most-covered dimensions recover worst — IIC 0.61, L 0.67, SM 0.73 — exactly the identifiability worry the coverage distribution predicted: a dimension active on every item is rarely isolated by the data. Raising τ fixes it (at τ = 1: 0.83 / 0.90 / 0.92) at the cost of Instrumental Reasoning (0.89 → 0.76). τ = 1 is adopted because it maximizes the worst dimension, not the mean.

Intercept. Because pooling depends only on the margins c_k − λD_jk, raising every capability and lowering the intercept by the same amount is observationally identical. A free per-agent intercept therefore leaves overall level unidentified (level recovery r = 0.12) while shape survives (r = 0.92); a single intercept shared across the catalogue lifts level recovery to 0.98, effective posterior dimensions from 2.7 to 6.7, and downstream suitability recovery from ρ = 0.14 to 0.91 — essentially matching an oracle that fixes α to its simulated truth.

The consequence is a scope limit on every level number on this page. Capability levels are identified only relative to the other agents in the catalogue; adding or removing a system shifts the common reference point, and absolute levels must not be compared across separately fitted catalogues. Profile shape is identified regardless. This is the same shape of caveat Item Response Theory for LLM Benchmarks carries about population-relative difficulty, relocated: there the items are anchored to a model population, here the agents are.

And the recovery analysis proves less than it looks like it does, which the authors say outright: the same modelling parameters generate the simulated data and fit it, so "the analysis establishes internal recoverability, but it does not establish the empirical validity of those modelling assumptions." Whether real task performance actually pools soft-min at τ = 1 is untested.

Stage 2: 410 workers spend 100 points#

18 work activities adapted from O*NET work-activity categories (Table B1, verified), chosen to be commensurable across six job domains standing in for the departments of a product company: Warehouse or logistics (WL), Manufacture/maintenance/repair (MMR), Numerical/data/programming (NDP), Administration/organisational/planning (AOP), Customer service/marketing/HR (CMH), Hospitality/sales/client care (HSC).

Sample. 125 respondents recruited through collaborating companies plus 414 through Prolific; 539 total, 410 after quality control (completion, ≥ 50% on a capability quiz, ≥ 10 minutes — removing 59, 66 and 4 respectively). Table B2 is verified and its margins close exactly in both directions (85 + 325 = 410; 125 + 414 = 539). Mean age 40.3, 49.8% female, typically degree-holding, 5.7 years in the current role against 11.8 in the wider field, AI attitudes mildly positive (mean ≈ 58/100). Nationality is not reported in the raw beyond the recruitment platforms; the companies are unnamed.

Instrument. Four stages, ~30 minutes: demographics → pick the five most important activities, rank them, and report hours/week on each → capability familiarisation (all 18, each with definition, cross-domain examples and an illustrative image) plus a matching quiz doubling as an attention check → for each selected activity, pick the five most essential capabilities and distribute 100 points across them, framed as equipping a "robot helper."

Aggregation is frequency-adjusted: unselected items score zero rather than missing, so a capability's weight encodes both how often it is chosen and how heavily it is loaded when chosen.

Sample validation. Recruitment sources are very unevenly distributed across domains — the manual-physical domains are almost entirely online (WL 36 of 37 post-QC, MMR 33 of 33) — so company and online importance matrices are compared before pooling. Agreement by activity is cosine 0.91 / Pearson 0.80; by capability 0.85 / 0.52, the lower figure driven by rarely selected capabilities estimated from few observations. A split-half noise ceiling puts within-source reliability up to 0.94 and the disattenuated between-source r at ≈ 1 — the two samples agree about as well as two random halves of one sample. That is a genuine reliability result, and it is worth being precise that it is only that: it shows the instrument is self-consistent across recruitment channels, not that what it elicits is a valid account of cognitive demand.

What workers say matters#

Activities (frequency-adjusted importance I_t, averaged across domains): Problem solving 2.04, Decision making 1.78, Checking 1.44, Researching 1.35, Computer use 0.96 at the top; Listening 0.29, Coding 0.38, Managing resources 0.39, Data manipulation 0.39 at the bottom. Because the weighting multiplies perceived importance by selection frequency, broadly applicable activities outrank ones intensely important to a minority — Coding's 0.38 is a base-rate artifact, not a claim that coding is unimportant where it happens.

Problem solving and Decision making rank at or near the top in every domain; the domains separate on supporting activities. WL and MMR put Checking at 2.43 and 2.45 against 1.44 overall and Tool use at 1.32 and 1.88 against 0.67. NDP elevates Computer use and Analysing data to 1.89 each, with Coding 0.68 and Data manipulation 1.06. AOP uniquely raises Long-term planning (0.94 vs 0.47). HSC peaks on Building rapport at 1.83.

Importance is not time. Checking ranks near the top on importance while occupying 10.2 hrs/wk; Computer use consumes the most time at 15.7 hrs/wk on only moderate importance, with Tool use 14.7, Managing people 13.4 and Admin 10.8 likewise over-indexed on hours. Two different deployment cases follow — augmenting high-value work versus draining high-volume work — and the paper carries both (Figures C6 and C7).

Capabilities. The task × capability matrix is dominated by a shared cognitive core: Planning, Semantic Memory, Working Memory, Language and Procedural Memory take the greatest weight across nearly every activity. Activities differentiate on secondary capabilities — interpersonal work on social cognition, analytical work on pattern recognition, creative work on mental simulation. Domain-specific matrices correlate cell-for-cell at r = 0.53–0.77 (mean 0.63), with the largest departures being WL +2.1 on Spatial Reasoning & Navigation, MMR +2.5 on Planning, and HSC +1.7 on Theory of Mind and +1.2 on Emotion Perception & Empathy — all secondary, none touching the core.

The wiki has been circling this convergence from the other side. Task Crossover finds 43.5% of occupation-specific AI use is another occupation's work, and Task Saturation: Broad but Shallow AI Diffusion finds AI reaches a median 21% of a reached occupation's tasks. A stable cross-occupational cognitive core is a mechanism that would produce both: if occupations differ mainly in secondary capabilities, the tasks that cross occupational boundaries are the ones loading on the core.

What the six systems look like#

Six systems from two developer families — Gemini 2.5 Flash, Gemini 3 Flash, Gemini 3.1 Pro, GPT-4o-mini, GPT-5-nano, o4-mini — fit jointly with a shared intercept, 2,000 draws/chain × 4 chains, R̂max ≤ 1.013 and bulk ESSmin ≳ 350 (Table C5, verified).

SystemBattery accuracyOverall level c̄
Gemini 3.1 Pro73.3%3.70
Gemini 3 Flash68.4%3.38
o4-mini61.8%2.88
GPT-4o-mini48.7%2.79
GPT-5-nano56.0%2.75
Gemini 2.5 Flash44.8%2.58

Note that accuracy and inferred level already disagree: GPT-4o-mini is last on raw accuracy (48.7%) and fourth on capability (2.79), ahead of GPT-5-nano, which answered 7.3 points more items correctly. That is the IRT reordering effect Item Response Theory for LLM Benchmarks documents, reproduced here with rubric-derived rather than population-derived difficulty — and the mechanism is visible in the per-dimension table: GPT-4o-mini's correct answers concentrate on high-demand Language and Social Cognition items.

The per-dimension picture (Table C6, verified cell-for-cell)#

SystemIICLSMAPSIREMSCOP
Gemini 3.1 Pro4.905.036.254.251.442.674.790.27
Gemini 3 Flash5.064.966.122.371.322.934.180.11
o4-mini2.922.906.341.631.713.234.42−0.10
GPT-4o-mini2.214.935.730.032.052.984.180.19
GPT-5-nano4.192.226.171.341.304.512.59−0.34
Gemini 2.5 Flash2.914.072.952.31−0.513.064.321.58
Mean3.704.025.591.991.223.234.080.29

(Posterior SDs, omitted here, run 0.12–0.52; the mean row reproduces Table 6 exactly, an independent consistency check across two tables on different pages of the paper.)

The ordering of dimensions is nearly invariant across systems. Semantic Memory is the top dimension for five of six (c ≈ 5.7–6.3); the exception is Gemini 2.5 Flash, which peaks on Social Cognition. Action Planning & Simulation, Instrumental Reasoning and Object Permanence are the bottom three throughout. Current systems are strongest on knowledge, language and social understanding, and weakest on planning and reasoning about actions and objects.

What separates the leaders is not knowledge. The Gemini 3 models' largest advantage is Information Integration & Control (5.06 and 4.90 against c ≈ 2–4 for the rest) and Action Planning & Simulation, where Gemini 3.1 Pro's 4.25 is 1.88 above the next-highest system's 2.37 — the single largest gap anywhere in the table. Knowledge- and language-related capabilities differ comparatively little. The reading the authors take: recent progress toward agentic behaviour has come from higher-level coordination, not from robust reasoning about objects and their interactions — Object Permanence stays at 0.27 and 0.11 for the two leaders, and Instrumental Reasoning at 1.44 and 1.32.

The jagged profiles are individually informative. Gemini 2.5 Flash is the only system with Object Permanence above 1 (1.58) and the only one with a negative Instrumental Reasoning (−0.51), while its Semantic Memory (2.95) sits 2.8–3.4 below every other system. GPT-5-nano is the Episodic Memory outlier (4.51) and simultaneously the Language floor (2.22) and Social Cognition floor (2.59). GPT-4o-mini has the highest Instrumental Reasoning (2.05) and an Action Planning level of 0.03 — effectively none. These are not rank-ordered systems on one axis; they are different shapes, which is the paper's point.

The headline claim, and exactly what carries it#

"AI systems differed more across cognitive dimensions than across model families." Quantitatively, what that rests on is the comparison of two spreads:

  • Across dimensions (Table 6 means): 0.29 (OP) to 5.59 (SM) — a range of 5.30.
  • Across systems (Table C5 overall levels): 2.58 to 3.70 — a range of 1.12.

Roughly 4.7×. No variance decomposition, ANOVA, or formal effect-size comparison is reported anywhere in the paper — the claim is made by pointing at a table and a radar plot (Figure 3), and it should be quoted as the descriptive observation it is. It is also partly a construction: the dimension spread is measured on a scale whose dimension-level anchoring is explicitly not absolute (a level-4 Working Memory demand and a level-4 Theory of Mind demand are not equated), so comparing 0.29 to 5.59 across dimensions is comparing numbers the paper elsewhere declines to equate. The comparison of system levels is the better-founded of the two and it is the smaller number.

The defensible version of the finding, which the per-system table does support, is about shape: six systems from two labs across three generations produce profiles whose rank ordering of dimensions is nearly identical, and family membership predicts profile shape poorly (Gemini 2.5 Flash looks less like Gemini 3 Flash than o4-mini does on five of eight dimensions).

A vintage caveat the paper does not raise#

The catalogue is not a frontier catalogue: five of the six systems are explicitly the cheap tier (two Flash models, two minis, one nano), with Gemini 3.1 Pro the only full-size system and no Claude, Grok, or open-weight model anywhere. Because the intercept is shared within the catalogue, every level on this page is anchored to a mostly-small-model reference point, and the paper's own scope rule forbids carrying those levels to a differently constituted catalogue. The "six AI systems" of the abstract is therefore six systems of a particular and unrepresentative kind.

Stage 3: suitability, and how little it moves#

Suitability S_at is the importance-weighted power mean of order p over ratio-scale capabilities, with p = 0 (weighted geometric mean, the natural midpoint for ratio-scale quantities) and sharpening exponent s = 1 as neutral defaults. Both are decision-stage policy parameters, not inferred from data: p governs substitutability between capabilities (p → −∞ is weakest-link, p > 1 lets one standout strength dominate), s governs how concentrated the weights are. Capability posteriors are pushed through independently and weight uncertainty is propagated by sampling from a Dirichlet centred on the elicited profile, so suitability is itself a posterior.

Across the 18 activities: Gemini 3.1 Pro is the most suitable system for every one (log S ≈ 4.1–4.6), Gemini 3 Flash second throughout (≈ 3.3–4.1), and the remaining four form a lower, heavily overlapping group (≈ 1.7–3.4). The ranking barely changes from activity to activity — "systems differ from one another far more than activities differentiate the systems."

That stability is a direct consequence of the shared cognitive core. Almost every activity loads on knowledge, language, planning and cognitive control; systems differ little on the first two and a lot on the last two; so the same ordering falls out of almost any weighting. The exceptions are the socially oriented activities — Building rapport, Listening, Communicating — which load Social Cognition and Language hard enough to partially break the ordering, and where GPT-4o mini rises to the top of the mid-pack (log S ≈ 3.4).

Uncertainty tracks sample size in the obvious way: Coding's importance profile rests on ~20 respondents and has the widest interval on the chart; Problem solving, well-sampled, the narrowest.

Deployment priority. Suitability says a system can do a task, not that doing it matters. Multiplying by task importance (P_at = I_t · S_at) reorders toward Researching, Admin and Analysing data as the highest-priority targets overall. The single-company case study (Company X, 35 post-QC responses, predominantly AOP and CMH) surfaces Communicating and Researching as its clearest opportunities, with Problem solving and Computer use as high-importance capability gaps.

Sensitivity. Sweeping p and s broadly, nearly 12 of the 15 pairwise system orderings survive on average, each task shows only three to seven distinct rankings across the full compensatory sweep, and Gemini 3.1 Pro stays top for 17 of 18 activities (Admin flips only at the most extreme compensatory setting p = 2). Reordering, where it happens, is confined to adjacent systems whose profiles already overlap.

What it does not measure#

This is the section to read before quoting a suitability number anywhere.

It measures importance, not demand — so nothing here is calibrated. The authors state it plainly: the pipeline "measures task importance not task demand," so suitability scores "indicate which agents are better matched to a task, but not the probability that an agent will successfully perform it." The capability half is on a demand-anchored scale; the requirements half is a relative-importance budget. The stated fix — extending rubric-based demand annotation to workplace task instances — is judged infeasible at organisational scale and left undone.

Nothing is validated against an observed outcome. The two validations performed are (a) synthetic-agent recovery, which the authors say establishes internal recoverability only because the same model both simulates and fits, and (b) company-versus-online questionnaire agreement, which is a reliability check on the instrument. There is no deployment outcome, no productivity measure, no realized adoption series anywhere in the paper, and no correlation reported against GDPval Benchmark, the Anthropic Economic Index, ATLAS or any other external exposure measure. The framework's whole selling point — that it predicts where systems will and won't work — is untested.

The battery under-samples exactly where the systems look weakest. Assembled primarily from text-based evaluations, it under-represents multimodal perception, long-horizon planning and interactive tool use. The authors concede "true differences between systems on those dimensions may therefore be larger than the present profiles indicate" — which means the low Object Permanence, Instrumental Reasoning and Action Planning floors are a joint property of the models and of a text battery, and cannot be attributed to the models alone. Table 2's observation that no dimension reaches demand level 5 on any item makes this concrete: the battery's hardest items do not reach the top of its own scale.

These are base models, not deployed systems. Scaffolding is precisely what compensates for the weaknesses measured here: "a monitor that tracks background state effectively supplies object permanence, while well-specified tooling reduces the demand on affordance perception by making available actions explicit." A capability profile of a bare model is not a profile of the agent anyone deploys — the same gap Harness Shrinkage as Models Improve treats from the harness side.

Human-derived requirements may mischaracterise what an AI uses. The paper's own example: human programmers lean on planning and procedural memory, while contemporary systems often succeed on statistical pattern matching, so a requirements profile elicited from humans can weight the wrong capabilities for a non-human agent. The authors say this "cannot be resolved from within the framework itself" and requires validation against deployment outcomes — the thing the paper does not have.

Introspection bounds the requirements side. Metacognition, spatial reasoning and object permanence were selected rarely, which the authors attribute to low salience to conscious reflection rather than low contribution, compounded by the five-capabilities-per-activity cap. Automatic cognitive processes are exactly the ones self-report misses — and, awkwardly, object permanence is simultaneously the dimension on which every profiled system scores lowest and one workers almost never nominate, so the two halves of the pipeline are both blind in the same place.

One circularity the paper discusses, and one it does not. It addresses the general worry — an LLM annotating the demands of items on which LLMs are evaluated — by separating roles (the annotator is a rubric-applying classifier, not a subject) and noting the rubrics encode human expert judgement. What it does not note is the specific overlap in this experiment: Gemini 3 Flash is both one of the two demand annotators and one of the six profiled systems. Its own battery difficulties are half-authored by itself. The consensus averaging with GPT-4o dilutes this and the demand scale is defined before any system is run, so it is not obviously fatal — but it is unaddressed, and an annotator/subject disjointness check is a cheap thing the next version should run.

Where this sits among exposure instruments#

Read as a labor instrument rather than an evaluation method, this is an eighth instrument on Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated's list of seven, and it is the first one whose primitive is neither a task description nor a usage log. Every instrument on that page scores work — Eloundou's GPT-4 task ratings, Felten's 52 O*NET abilities, Webb's patent verb-object pairs, Frey's whole-job automation probabilities, the Anthropic/OpenAI usage-derived ones — and asks how much of it AI reaches. This one scores the model on cognitive constructs, scores the work on the same constructs, and derives exposure as the match.

Felten's ability-based instrument is the nearest relative and the contrast is instructive: Felten rates O*NET's 52 human abilities for "AI suitability" using 2,000 MTurk workers, which is a crowd's forecast about technology. This rates the capabilities of a named system from 19,535 measured item outcomes, which is not a forecast about anything. Steele & Cruz's finding that exposure instruments correlate when they share a data source, not when they claim the same construct makes this one's placement a genuinely open empirical question — its data source (benchmark response matrices) is shared with no other instrument on the list, which by that page's generalization predicts it will correlate with none of them.

Nobody has run that comparison. It is cheap to run — the profiles and importance matrices are published (github.com/Kinds-of-Intelligence-CFI/Task-Suitability-Profiles) and O*NET crosswalks exist for the 18 activities — and it is the single most informative thing anyone could do with this paper.

Connections#

  • Item Response Theory for LLM Benchmarks — the same measurement family, with the difficulty parameter sourced the opposite way. ATLAS estimates item difficulty from ~4,000 models' responses, making b population-relative and subject to temporal drift; this annotates demand from expert rubrics applied by an LLM judge, so the bank never goes stale relative to a model population. Both end up with a scale anchored to something contingent — there the examinee population, here the fitted catalogue via the shared intercept — and the honest summary is that neither has produced an absolute capability scale, only two different relativities. Both also reproduce the accuracy-reordering effect: GPT-4o-mini is last on raw battery accuracy and fourth on inferred capability here
  • Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — the eighth instrument, on a construct axis none of the seven use, and the first whose primitive is a measured capability rather than a rated task. Its predicted correlation with the other seven, under that page's data-source generalization, is near zero — untested
  • Benchmark Score Redundancy — the collinearity problem arriving one level down. That page's matrix is models × benchmark scores at effective rank 2; here the collinearity is among demand columns of 19,535 items, forcing 16 capabilities down to 8 composites by clustering at 1 − r = 0.5. Same remedy, different granularity — and a useful third data point against CollabEval's finding that item-level matrices need ~16 components: the demands annotated on items are collinear even where the response patterns are not
  • Economic Benchmark Construct Validity — the one-capability-or-many question with the axes swapped. Zhu factor-analyses models × benchmarks and finds one factor at 74.5% of common variance that tracks release date at R² = 0.505; this paper pre-specifies the factors from cognitive theory and measures items against them, so its dimensionality is a design choice rather than a finding. Zhu's calendar confound is what this design is built to avoid — a rubric demand level does not move with release date — but the shared intercept reintroduces a version of it, since the catalogue's reference point shifts whenever its membership does
  • GDPval Benchmark — the outcome-measured counterpart, and the comparison this paper never makes. GDPval asks whether a model's deliverable beats a 14-year professional's on real work product, scored as a blinded pairwise win rate; this asks whether a model's cognitive profile matches what workers say their activities require. One measures output against the incumbent and has no capability decomposition; the other decomposes capability and never observes an output. Running the six profiled systems on GDPval's open gold subset and correlating with suitability is the obvious unbuilt validation
  • Measuring Beyond Accuracy Saturation — the third answer to "an aggregate score tells you little." That page adds axes to a saturated benchmark (reliability, cost, scaffold contribution, human uplift); Item Response Theory for LLM Benchmarks changes the scale; this one changes the unit of measurement — it reorganises existing benchmarks around the cognitive constructs items recruit rather than the task domains they nominally belong to, which needs no new instrumentation and no new items, only a rubric pass over an existing bank
  • Task Crossover — the stable cross-occupational cognitive core is a candidate mechanism for crossover. If six occupational domains' capability matrices correlate at r = 0.53–0.77 and differ only in secondary capabilities, tasks loading on the shared core are precisely the ones that should cross occupational boundaries — which is what 43.5% of occupation-specific use being another occupation's work looks like from the cognitive side. Also the shared dependency: both take their work-activity vocabulary from O*NET, so both inherit whatever that taxonomy gets wrong
  • Task Saturation: Broad but Shallow AI Diffusion — the depth counterpart. ATLAS finds AI reaches a median 21% of a reached occupation's tasks; a shared cognitive core plus system profiles that are strong on knowledge/language and weak on planning/objects predicts exactly that shape — broad reach through the core, shallow penetration wherever an activity's secondary demands fall on APS, IR or OP
  • Machine Self-Report Psychometrics — psychometrics for models, pointed the other way. That page's finding is that human questionnaires fail when aimed at models and needed a purpose-built instrument; this one refuses to administer human instruments at all, and builds the capability space from item demands instead. Both converge on the same diagnosis — anthropomorphic transfer is the failure mode — and pick opposite escapes: rebuild the instrument, or rebuild the scale the existing instruments are scored on
  • Harness Shrinkage as Models Improve — why a base-model capability profile is not a deployed-agent profile. The paper's own examples are a state monitor supplying object permanence and explicit tooling reducing affordance-perception demand; every dimension where these systems score lowest is a dimension a harness is built to cover
  • Returns to Expertise in Agentic Coding — the requirements side's blind spot, named from the other direction. The questionnaire elicits the capabilities an activity needs from any agent, which is deliberately the part of work that does not differentiate one worker from another; domain expertise, the thing that most amplifies an agent, is excluded from the capability set by construction
  • Benchmark Convergent and Discriminant Validity — the collinearity this page finds in item demands, found again in benchmark outcomes and at the level of labels. Desai et al. show reasoning, knowledge and comprehension benchmarks ranking 53 models as alike across labels as within them (within − between ρ = −0.00), just as this page's 16 annotated capabilities collapse to eight composites because hard items recruit many at once. Neither result says the capabilities are one thing. Both say the instruments cannot tell them apart. Their safety grid is the counterpoint: same-label safety benchmarks barely converge, so "many dimensions" is the better description there

Open Questions#

  • Does capability-profile suitability correlate with any realized outcome? Nothing here is scored against a deployment, a productivity measure, or an existing exposure instrument. Directly falsifiable and cheap, because the profiles and importance matrices are published: correlate the six systems' per-activity suitability against GDPval win rates on matched occupations, and the 18 activities' suitability against the task-level observed-exposure measures that already exist. A near-zero correlation with the task-rated instruments would confirm Steele & Cruz's data-source generalization; a strong one would make this the first instrument to bridge families.
  • Are the Object Permanence, Instrumental Reasoning and Action Planning floors a property of the models or of a text-only battery? The paper says the battery under-samples exactly these and that true differences "may be larger," and Table 2 shows no dimension reaches demand level 5 on any item — so the floors are measured against a ceiling the battery itself sets. Falsifiable by re-running the pipeline with multimodal and agentic benchmarks annotated under the same rubrics and checking whether the bottom three dimensions separate the systems more, less, or in a different order.
  • Does an annotator that is also a subject bias its own profile? Gemini 3 Flash annotated the demand levels of the battery on which Gemini 3 Flash was then profiled, and the paper's circularity discussion does not reach this specific overlap. Falsifiable with the released annotations: re-fit all six systems using GPT-4o's demand matrix alone, then Gemini 3 Flash's alone, and check whether either Gemini 3 model's inferred profile moves more under its own annotator's matrix than the OpenAI systems do.

Sources#

  • What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks — Desai et al., arXiv 2609.08812, 2026-09-08, COLM 2026, empirical. Cited here only in the connection above: §4.1–4.2 and Table D.4 (capability within − between ρ −0.00; safety non-convergence). Full treatment on Benchmark Convergent and Discriminant Validity

  • Using profiles of cognitive capability to assess AI suitability for workplace tasks — Jonathan Prunty, Marko Tešić, Patrick Quinn, José Hernández-Orallo & Lucy Cheke (University of Cambridge / UK Department for Science, Innovation and Technology / Universitat Politècnica de València), Using profiles of cognitive capability to assess AI suitability for workplace tasks, arXiv 2608.25623, 2026-08-26, 44pp, technical report, empirical — tier kept. Code, rubrics and benchmark annotations released at github.com/Kinds-of-Intelligence-CFI/Task-Suitability-Profiles. Cited here for §3.1 (capabilities, rubrics, battery construction, inter-rater screen, clustering), §3.1.3 and §A.2–A.3 (Measurement Layouts, margins, soft-min pooling, intercept identifiability), §3.2 (questionnaire instrument and aggregation), §3.3 (power-mean suitability), §4.1–4.4 (recovery, profiles, task importance, mapping, sensitivity, Company X) and §5.1 (limitations). COI: the authors acknowledge the support of Accenture, and one co-author (Quinn) is at the UK government department that also commissions AI-adoption research cited in the paper's own motivation; participating companies are unnamed. No competing-interests statement appears. Parse notes. The raw is docling-derived (2.126.0 / docling-mlx 0.1.1, confidence_grade: excellent) and carries zero [!note]/[!warning] repair blocks, so every table was treated as unrepaired. Tables 1, 2, 4, 5, 6, A2, B2, C5 and C6 were reconciled cell-for-cell against pdftotext -layout on (pp. 5, 6, 10, 11, 12, 22, 30, 33) and are byte-exact; Table 1's 22 item counts additionally sum to its own printed total of 19,576, and Table B2's margins close in both directions. Tables A1, A3 and B1 are definitional text tables, cross-checked against the prose. Tables C9 and C10 are badly shifted and welded in the docling parse — values slide into neighbouring cells (row 1 of C9 reads 3.64 [3.41, 3.87] 3.16 where the PDF has a clean two-line cell) and from row 2 onward the mean/CI pairs split across rows. Both were verified as damaged against page 40 and no number from either is cited on this page or anywhere in the wiki; the domain-level suitability and deployment-priority figures quoted here come from the §4.4 prose only. Figures viewed under the image two-pass rule (figure↔image mapping confirmed sequential, image_000000–image_000006 = Figures 1, 2, 3, 4, 5, 6, B1): Figure 3's radar plots confirm the Table C6 shape ordering and the Gemini 3.1 Pro APS spike; Figure 5's forest plot confirms the log S ≈ 4.1–4.6 / 3.3–4.1 / 1.7–3.4 banding, GPT-4o mini's rise on Building rapport, and Coding as the widest interval. Figure 4 (the 18 × 16 task-capability heatmap) was viewed but is not legible enough at the delivered resolution to read individual cells, and no number is taken from it — the shared-core finding on this page comes from the §4.3.4 prose. One in-paper arithmetic note, reconciled rather than flagged: Table A2's note reports 19,531 items rated validly by both annotators against the 19,535-item battery, and the paper explains the four-item difference itself (the battery requires valid ratings only on the 16 retained capabilities).

§ end
Cited by 14
Related articles