H
Howardism
Plate IIAI Economics & Labor中文HOWARDISM

AI Adoption in Scientific Work

Three instrument families on one profession. Google ATLAS telemetry + survey: scientists are the economy's heaviest AI users (SOC 19 over-indexes 2.7x; 46.6% use AI daily), LLMs and 2,690 specialized models are complements (elasticity 0.6->0.2), and a self-reported 6.9 hours/week saved does not become discovery because the bottleneck moves downstream (43.5% say so; 45.7% spend over a quarter of it verifying; 48.8% tilt to safer questions). Against it, a latent class analysis of 3,785 PhD students finds attitudes arranged by task: 51.5% comfortable with AI summarising literature against ~30% for writing, analysis and experiment design, in a 44% "division of labour" profile. Beside both, publication traces (~22% of computer-science output carrying LLM-modified text by September 2024) - blind to analysis, but on writing they find behaviour where the survey finds refusal

Article metadata
Publication details
Published:September 22, 2026
Filed:Concept
Domain:AI Economics & Labor
Reading:38 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for AI Adoption in Scientific Work

Sources#

Summary#

AI in Science: Early Insights (Codreanu, Imas, Mateos-Garcia et al., Google + Google DeepMind + MIT FutureTech, September 2026) is the first attempt to measure AI in science with a usage-telemetry instrument rather than publication traces. Its four findings, in the authors' order: scientists lead other occupations in AI use; general LLMs and specialized scientific models are complements, not substitutes; scientists report large time savings (just below 7 hours/week); and those savings are not becoming discoveries, because the binding constraint moved downstream into physical experimentation, clinical validation, and verification of AI output.

A second instrument now sits on this page (2026-09-23). Angelini & Lyrvall's latent class analysis of 3,785 PhD students in Nature's Graduate Survey 2025 measures nothing ATLAS measures — not usage, not time saved, not output — and asks instead which research tasks researchers think AI belongs in. Its answer is the normative complement to the one above: acceptance in science is organised by task, and the boundary falls between literature work and the activities that carry authorship and epistemic responsibility. The two instruments also disagree in a way neither can settle — see the contradiction below.

The last finding is the one that earns this page. It is the complements thesis measured inside the research production function, and it is the adopter-side counterpart to the capability claims — the same verification gap those pages argue about theoretically, reported as a time budget by 637 working scientists.

Evidence note. empirical — three measured data sources with documented pipelines, published classifier validation, and author-stated limitations. Two COI facts attach to every number. The vendor measures its own product: the telemetry is Gemini traffic, analyzed by Google, using Google's own classifiers, in a paper arguing that AI raises scientific productivity. The survey was commissioned by Google DeepMind and fielded by a third party (More in Common); third-party fielding does not make the sponsor's interest go away. Everything in the survey is self-report — time saved, bottleneck location, verification share, output change — from a screened non-probability panel with no population frame, reported unweighted with no margin of error, and the authors themselves name selection toward AI-enthusiastic respondents as a reason the adoption rate and the time savings may be overestimates. The paper establishes associations, not causality, and says so.

Not a second data window#

Worth stating first, because the follow-on format invites the opposite assumption. The telemetry here is the same ATLAS v1.0 corpus — "approximately 15 million anonymized interactions across the Google Gemini App, Google AI Mode, and API interfaces, sampled in early April 2026" (§2.2.1), i.e. the April 6–19 window behind Task Saturation: Broad but Shallow AI Diffusion. What is new is the filter and the taxonomy, not the collection: a three-stage pipeline (work/non-work classifier → six research-intensive 2-digit SOC groups → a bespoke science classifier) narrows 15M interactions to about 360,000 science interactions, which are then mapped onto OpenAlex disciplines and a new MIT FutureTech task taxonomy instead of O*NET.

So this report adds no time dimension to the ATLAS program. The enterprise and agentic exclusions also carry over verbatim: no paid Gemini API, no enterprise contracts, and — stated as a limitation in §6.1 — no agentic tools such as Antigravity, "incorporating agentic data into our analysis is an important next step."

Scientists are the heaviest users, and the usage is more sophisticated#

  • Over-representation. Gemini volume across the six research-intensive SOC families runs 1.8× their share of US employment; for SOC 19 (Life, Physical, and Social Science) it is 2.7×, and for the core STEM detailed occupations inside SOC 15/17/19 it is 5.8× (§3.1 and fn. 7). The paper checks itself against the rival instrument: the Anthropic Economic Index finds SOC 19 over-represented about 4.5× in the US as of June 2026 (fn. 8) — same direction, different level, which is the standing pattern between these two programs.
  • Breadth. Usage clears the 25-user privacy threshold in 195 of 217 scientific subfields (§3.2). Physical Sciences including Computer Science take 55.3% of science interactions, Social Sciences 21.4%, Health Sciences 14.1%, Life Sciences 9.2%; Computer Science alone is nearly three in ten.
  • Scaling. Within fields, a 1pp increase in a field's share of OpenAlex researchers predicts ~0.7pp more share of science interactions, explaining just under 30% of the variance — so most field-level variation is not headcount. Across countries the same regression gives ~0.9pp and explains about three quarters of the variance (§3.2).
  • The interactions themselves are heavier. Against the average work conversation, science interactions run +19% tokens, +11% turns, +7% multimodal, and +26% on ATLAS's domain-expertise score (§3.1, conversational surfaces only) — the cross-lab echo of compute tracking the value of the work.

The task mix (Figure 4)#

Level-1 shares of science Gemini interactions, weighted:

Task areaShare
Analyze and model quantitative research data42.5%
Communicate research findings and stakeholder information14.0%
Develop products, prototypes, and process technologies11.4%
Conduct experimental and sample processing operations9.8%
Teach and train students and research staff6.4%
Manage research projects and operational processes3.9%
Conceptualise, design, and propose research3.0%
Coordinate clinical study execution and participant operations2.8%
Develop and troubleshoot research assays and software2.3%
Manage research data, integrity, and quality1.5%
Ensure quality, compliance, and regulatory documentation1.3%
Collect research data and participant information1.1%

The authors flag the communication share as a lower bound: their pipeline over-predicts "analyze and model" when a prompt carries an attachment, and routine writing is more likely to be filtered out as generic text editing (fn. 16). The field-level over/under-representation cuts (Table 1) are taken here from §3.3's prose rather than the table, which parsed with its label column split (see Sources).

LLMs and specialized models are complements#

The report's cleanest measured result, and the one that does not depend on self-report.

The inventory: 2,690 specialized scientific models published since 2012 with a paper and an official code repository, drawn from a larger 5,501-model collection (4,858 after deduplication) assembled by agentic web search over repositories and literature, merged with Epoch AI's database (August 2026 snapshot) and enriched from OpenAlex. They cover all 26 OpenAlex fields and 83% of subfields; OCTO extracted 4,475 downstream tasks from their abstracts, covering 17.6% of the taxonomy's Level-3 tasks, or 24.3% excluding operational and teaching tasks (§4.2).

Both families look alike at the top of the taxonomy — quantitative analysis and modeling dominates each — and diverge completely underneath. The log-log elasticity of specialized-model task share on Gemini task share falls monotonically with resolution: just above 0.6 at Level 1, 0.5 at Level 2, 0.2 at Level 3 (and slightly negative at Level 3 if tasks present in only one corpus are included). Specialized models concentrate in disease and clinical-outcome prediction, molecular construct engineering, and simulation; Gemini absorbs statistical analysis, code troubleshooting, literature synthesis, and drafting (§4.3). The authors caution that part of the decline is attenuation bias — classifier accuracy falls as granularity rises — so the nominal elasticities are upper bounds on the overlap.

Two supporting facts, each with its own caveat. The specialized models are highly cited — 1.3M citations since 2012, 49% of the models in the top 1% of normalized citations in their field and 83% in the top 10% — but the authors note the sampling frame selects for notability, so this is qualitative evidence of reach rather than a measured citation premium (fn. 22). And development is concentrated while use is not: the top 10 countries hold 84% of the models, LMICs barely appear as developers, yet China accounts for ~40% of citations and LMICs across Latin America, South-East Asia and Africa cite far more than they build — the paper's reading is that open models diffuse where the capacity to build them does not (§4.5).

The survey: 637 scientists, and what they say the dividend buys#

Fielded by More in Common between 27 July and 11 August 2026, online, on specialist research panels, through four screens (main job in science/clinical/life/social research; research or R&D as a primary part of it; self-identification against the UK Science Council definition; market research, marketing and consultancy screened out). No incentive structure is disclosed beyond "recruited through specialist research panels." 637 active scientists: 379 US, 258 UK; Physical Sciences & Engineering incl. CS 230, Life Sciences 167, Health & Clinical 147, Social Sciences 93; 356 senior (PI, professor, lab director, industry R&D manager), 234 mid-career, 47 early-career; sample mean just below 13 years of research experience (Appendix 3).

  • Intensity (Figure 11). 46.6% (n=297) use AI daily, 30.6% weekly, 15.5% sometimes, 7.2% rarely or never — 77.2% weekly or more. Of self-reported AI working time, ~70% goes to general LLMs (41% general-purpose chat and document LLMs, the rest coding agents and assistants) and ~30% to specialized models.
  • The week it is spent against. Data collection and experimentation just below 8 hours, data analysis and interpretation about 7 hours (just under a fifth of the week, and the part AI touches most), writing ~5, methodology ~5, knowledge acquisition ~5, conceptualization plus funding/admin ~9 combined (§5.2).
  • Time saved. Just below three quarters report net weekly time saved against 6% reporting net time lost; the mean is 6.9 hours per week. Reinvestment: just below 30% into more research output, ~21% into physical lab execution and data collection, ~19% into harder problems, ~18% into work-life balance and fewer hours, 12% into teaching and admin.
  • Output. ~84% report a net increase in lab or professional outputs over three years, just below 47% an increase of 10% or more; about 9 in 10 expect increases over the next three years, with the share expecting 25%+ rising from ~7% to ~18% (Figure 12).

The bottleneck moved downstream, and verification eats the dividend#

This is the finding that keeps the page honest about the one above it.

  • Where the constraint is now. Physical experimentation and data collection is the largest current rate-limiting bottleneck at 24% overall and 30% in Physical Sciences & Engineering, then data analysis (~21%) and writing (~14%) (§5.3).
  • Which way it moved (Figure 13A). 43.5% say their primary bottleneck shifted downstream over two years, against 13.8% upstream — a 3.1× asymmetry — with 42.4% unchanged.
  • The backlog (Figure 13B). 40.5% report a larger backlog of untested hypotheses than three years ago, 24.8% smaller, 33.8% unchanged.
  • The verification tax (Figure 13C). Of scientists who save time with AI, 89.3% spend more than 10% of the saved time verifying, debugging or fact-checking output; 45.7% spend more than 25%; 9.3% spend more than half. The report says the tax is highest in the Life Sciences. Footnote 25 matters for how hard to lean on this: the downstream shift and the backlog growth are "statistically and economically very strongly correlated" with self-reported AI intensity, while the verification tax is only marginally significant and specification-dependent.

The paper's own summary of the mechanism is the O-ring one, citing Demirer et al. (2026), Gans & Goldfarb (2026) and Garicano et al. (2026): task-level acceleration does not propagate because "the residual tasks in the bundle of scientific work are harder to scale and absorb the time savings," so realizing the gains "may require redrawing job boundaries." That is the electrification argument with a wet lab in it.

The streetlight tilt#

The survey's most uncomfortable result, and the one with the least measurement behind it (Figure 14, all self-reported perceptions of the last two years):

PerceptionFavorableUnfavorable
Access to insights from other disciplines68.1%6.1%
Ambition of questions tackled67.3%6.4%
Breadth of research agendas65.0%7.1%
High-quality papers published59.8%9.1%
Less effort required per paper31.2%44.6%
Fewer low-quality papers published35.0%40.3%
Project risk: higher-risk vs safer27.5% higher-risk48.8% safer

Scientists report simultaneously that AI broadened their reach and pushed them toward safer, more incremental questions "where data and AI capabilities are well-established," by nearly 2:1. The authors name it a possible Streetlight Effect (Nagaraj & Tranchero; Hoelzemann et al.): AI lowers the cost of work in data-rich, benchmarkable problems, crowding out unstructured high-risk exploration where training data are scarce and physical validation is expensive. Note also that the same population reports both more high-quality papers (59.8%) and more low-quality ones (40.3% vs 35.0%) — "not just more output but different output," in the paper's phrase.

The normative side: 3,785 PhD students and a boundary that is not about capability#

Everything above measures what scientists do with AI and what the time buys them. Angelini & Lyrvall (Francesco Angelini, independent researcher, Italy; Johan Lyrvall, Inria Lille — arXiv 2608.25678, 2026-08-26, 15pp, empirical) measure the other half: which research tasks researchers think AI belongs in. The answer is not a level of acceptance but a shape — a boundary between literature work and the tasks that carry intellectual contribution.

The instrument. A secondary latent class analysis of Nature's Graduate Survey 2025 (Springer Nature with Thinks Insights & Strategy, fielded May–June 2025; the dataset is public on figshare): 3,785 self-selected PhD students across 107 countries, recruited through the Nature website, other Springer Nature digital products and targeted email. Five items, one per research activity — writing a research article, collecting and analysing data, designing experiments, tracking scientific literature, summarising scientific literature — each asking how comfortable the respondent is "with using AI tools (e.g. ChatGPT, Copilot and Claude)" for that activity; six ordered categories plus don't know, and no missing values on any item. Covariates are complete for 3,722 of the 3,785.

The population is narrower than "PhD researchers": 76% STEM, the remainder medical and health sciences, with no humanities and no non-medical social science at all. 38% are in their first two years, 87% full-time, 76% aged 34 or under, 54% male, and 15% come from a country on the UK's native-English-speaker exemption list. Self-selection is severe enough that the authors disclaim prevalence outright — the sample "does not support claims about the prevalence of these profiles among the wider population of PhD researchers." This is the same frame problem scientist surveys keep running into: there is no population register of PhD students to weight to.

The marginals already carry the finding (Table 1, n=3,785). Comfortable (very + somewhat) against uncomfortable (very + somewhat):

TaskComfortableUncomfortable
Summarising literature51.5%31.2%
Tracking literature43.1%35.0%
Writing a research article31.0%50.0%
Designing experiments30.5%45.2%
Collecting and analysing data29.5%49.9%

Exactly one task has a comfortable majority, and it is reading other people's papers. Don't know runs 4–8% and is highest on experiment design (8%).

Four profiles (Table 4; ICLbic selects four classes — 57,453 against 59,860 at three and 57,713 at five, Table 3).

  • 'Division of labour', 44% — the dominant profile and the paper's headline. It is never enthusiastic about anything: its very comfortable probability tops out at 0.10 (summarising). It is mildly positive on literature (0.50 somewhat comfortable on summarising, 0.41 on tracking) and mildly-to-firmly negative on writing (0.28 somewhat + 0.14 very uncomfortable), data collection and analysis (0.30 + 0.10) and experiment design (0.27 + 0.06). Its signature is the gradient, not the level.
  • 'Status quo', 34% — broadly refusing: very uncomfortable at 0.71 (writing), 0.76 (data), 0.71 (design) — but only 0.47 and 0.44 for tracking and summarising literature.
  • 'All-purpose', 16% — broadly comfortable: very comfortable at 0.51 / 0.51 / 0.47 on the three generative tasks, 0.65 and 0.76 on the two literature tasks.
  • 'Undecided', 7% — don't know at 0.57–0.81 across all five items. The authors explicitly warn against reading this as an attitudinal orientation; it may be non-use, unfamiliarity, or a response style. When these respondents do answer, they lean positive.

The structural point the abstract understates: the task ordering holds inside every class. Refusers are least uncomfortable about literature; enthusiasts are most enthusiastic about literature; even the undecided are least unsure about it. The literature-vs-contribution boundary is therefore not a property of the dominant profile — it is the axis the whole population is arranged along, and the four classes differ mainly in how far along it they sit and how sharply they bend at it.

Covariates (Table 5; two-step estimation, reference class 'division of labour', reference use category 'None'; n=3,722).

  • Field is the only structural predictor. STEM (vs medical and health sciences) is +0.307 (SE 0.11, p<0.01) toward 'status quo' — STEM students are the more refusing group, which is not the direction a "STEM is closer to the tools" prior would give. The other two contrasts are null (all-purpose 0.132, undecided −0.237).
  • Career stage and time commitment: nothing. First two years of the PhD gives −0.057 / 0.019 / −0.083, none significant; full-time likewise. The authors draw the inference that matters: research-usable AI has existed for roughly two years, so if these attitudes were formed by exposure-from-the-outset, the junior cohort would differ. It does not.
  • Use frequency moves people toward the differentiated profile, not the permissive one. Daily or weekly use (against none) gives −1.911 (SE 0.16, p<0.01) on 'status quo' and −2.402 (SE 0.24, p<0.01) on 'undecided', while 'all-purpose' sits at +0.434 (SE 0.23, not significant). Monthly/yearly use is negative and significant on all three (−1.128, −0.561, −1.987). So heavy users are overwhelmingly not refusers and not undecided, and are not measurably more likely to be across-the-board enthusiasts: use appears to resolve the boundary rather than dissolve it. Causation is unidentified — enthusiasm plausibly drives use as much as the reverse — and the design is cross-sectional.
  • Three coefficients the paper never discusses. Native English speaker: +1.047 (SE 0.14, p<0.01) toward 'status quo' and −0.777 (SE 0.26, p<0.01) away from 'all-purpose' — the largest effects in the table, and consistent with AI functioning as language support for L2 researchers rather than as a general research tool (the mechanism Hoomanfard & Shamsi document qualitatively for L2 dissertation writing). Male: −0.413 (SE 0.10, p<0.01) on 'status quo' and +0.272 (SE 0.12, p<0.05) on 'all-purpose' — women markedly more likely to be the refusing profile. Aged 34 or younger: +0.297 (SE 0.12, p<0.05) toward 'status quo' — the younger respondents are the more refusing ones. §4.1 reports only the field, career-stage and use-frequency rows; three of the paper's eight covariates carry significant coefficients that go unmentioned in the text.

What the instrument is not. The title says "normative boundaries"; the items ask about comfort. The authors concede the gap in §5.4: comfort "may reflect normative evaluations, but also familiarity with AI tools, access, confidence in using them, and beliefs about their reliability," so the profiles are "configurations of task-specific attitudes with a normative dimension, rather than a direct measures of shared scientific norms." They argue the normative reading from the shape — the boundary tracks authorship and epistemic responsibility rather than task difficulty — but the data cannot test it, and a plain reliability account predicts the same ordering: literature search and summarisation are the tasks current models are most dependable at. Nothing here measures task-level use, legitimacy judgments, disclosure behaviour, or outcomes. Use frequency enters only as an undifferentiated covariate.

The contradiction with the telemetry half of this page, which does not resolve cleanly. ATLAS's Gemini corpus puts 42.5% of science interactions in "analyze and model quantitative research data" — the single largest task area — while these PhD students report collecting and analysing data as the task they are least comfortable delegating (29.5% comfortable, 49.9% uncomfortable). Four readings are live, and the sources do not choose between them: (1) different populations — 637 senior-skewed US/UK working scientists against 3,785 internationally recruited PhD students; (2) different constructs — a share of interactions is not a rate of acceptance, and the comfort question is about delegating the activity, not about asking a model to check a regression; (3) a known classifier bias — ATLAS's own fn. 16 says the pipeline over-predicts "analyze and model" when a prompt carries an attachment and under-counts routine writing; (4) attitudes simply do not govern behaviour, which is the authors' own caution and the general result the instrument-aperture page keeps recording. The honest summary is that the two instruments are measuring different things about the same profession and agree on only one point — that the unit of analysis is the task, not the scientist. A third instrument family now narrows this (2026-09-23). Publication-trace scientometrics is blind to the disputed item — analysis leaves no trace in a paper — but on the adjacent item it separates the readings: writing is both the task this cohort most rejects (31.0/50.0) and the task where trace estimators find the most AI, at 22% of computer-science output in September 2024. See the trace-family section below; the effect is to make reading (4) the live one and to weaken (2) for writing.

Why it belongs next to the rest of this page. The governance argument in §5.1 is the process-versus-outcome one: if researchers attach normative weight to which activities were delegated, an evaluation regime that scores only output quality misses what they care about, and a blanket "AI was used" disclosure carries almost no information. Major publishers already distinguish AI use in conducting research, preparing a manuscript, and reviewing others' work, so the task-specific boundary is being written into policy independently of whether researchers hold it. §5.2 adds the doctoral-training version: the tasks in question are both production and the means by which research skill is acquired, so delegating them trades output for formation — the authors cite Bastani et al. (2025) and Fan et al. (2025), where AI assistance raised assisted performance and lowered unassisted performance afterwards. That is the agency-displacement construct arriving from a completely different method: resistance concentrates exactly where AI would take the generative part of the work.

A third instrument family: publication traces (2026-09-23)#

The two instruments above are telemetry (what people send a model) and self-report (what people say). Jedlička's survey (arXiv 2608.17970, 2026-08-18, practitioner-opinion) collects the third family in one place — traces left in the published literature. It measures none of it itself; §3 is a compilation of four indicators, and the compilation is the contribution, because the four are routinely quoted as if they counted the same thing. They do not.

Two indicators count AI as a topic:

  • Ding, Lawson & Shapira, Rise of Generative Artificial Intelligence in Science (Scientometrics 130 (2025): 5093–5114) — generative-AI diffusion across disciplines since 2022, "particularly pronounced" in the social sciences, psychology, education and the arts. The survey states the reason plainly and it is the whole problem with the indicator: those are the fields "where the societal implications of AI have themselves become major objects of investigation." A psychology paper about ChatGPT registers identically to a psychology paper written with ChatGPT.
  • Stanford's AI Index 2026 — AI-related publications in the natural sciences up more than a quarter between 2024 and 2025, to approximately 80,000 in 2025.

One indicator counts AI in the text:

  • Liang et al., Quantifying Large Language Model Usage in Scientific Papers (Nature Human Behaviour 9 (2025): 2599–2609) — a distributional estimate of LLM-modified text. As of September 2024, AI-assisted writing "may account for as much as 22 per cent of published work" in computer science, with statistics and physics lower and mathematics approaching 9 per cent. The survey's own caveat: "such estimates remain methodologically imperfect."

And one is the author's own, which is the weakest thing in the paper. Charts 1 and 2 are keyword counts from the nature.com search box (accessed 2026-06-11): across the Nature portfolio, papers referencing artificial intelligence rose more than seventy-fold from 2015 to 2025, surpassing 16,000 annually; in the flagship journal Nature, from ~50 papers/year in 2015 to more than 700 in 2025. There is no inclusion criterion, no disambiguation of the string, and one access date. Read from the chart images, the flagship series is not monotonic — it peaks near 247 in 2019, falls to ~172 in 2021, and only then turns into the post-ChatGPT climb (~213 in 2022, ~478 in 2023, ~572 in 2024, ~740 in 2025). The prose quotes only the endpoints and so reports a smooth exponential the author's own chart does not show. That dip is the tell that the series tracks editorial attention rather than research practice.

What the trace family can and cannot see#

It is the only one of the three that is unobtrusive, retrospective, and not owned by a vendor — no panel to recruit, no consent effect, no product to sell, and it reaches back to 2015 where the other two families have a single 2026 window each. What it cannot see is the thing this page is actually about: trace instruments observe published artifacts, not tasks. A paper is the end of the pipeline, so nothing upstream of writing leaves a trace at all — literature triage, a discarded analysis, an experiment that was designed and abandoned. And the topic indicators and the text indicator answer different questions that the survey runs together in one section.

What it does to the telemetry-vs-survey contradiction above#

It does not settle it, and the reason is worth stating precisely. The contradiction is about data analysis — 42.5% of ATLAS science interactions in "analyze and model quantitative research data" against 49.9% of PhD students uncomfortable delegating data collection and analysis. Analysis is invisible to a trace instrument: a regression run with a model's help and a regression run without leave the same paper behind. So the third family has nothing to say about the item in dispute.

Where it does bear is the adjacent item, and there it cuts one way. Writing a research article is the task the PhD cohort most rejects — 31.0% comfortable against 50.0% uncomfortable, the largest negative margin of the five — and writing is exactly the task where the trace instrument finds AI use at its most detectable, at 22% of computer-science output a full year before that survey was fielded. Of the four readings listed above, that is evidence for (4): stated attitudes do not govern behaviour, and it is the one reading the two existing instruments could not test, because both were pointed at the same moment from the same side of the respondent. It also weakens (2), the different-constructs reading, for writing specifically: "how comfortable are you with AI writing a research article" and "what fraction of published articles carry LLM-modified text" are close enough in referent that a 22%-versus-50%-uncomfortable gap is hard to attribute to construct drift alone.

Four things keep it from being decisive. The populations do not overlap (Liang's corpus is arXiv/bioRxiv/Nature-portfolio authorship; the survey is PhD students, only 76% STEM and with computer science not separately reported). "LLM-modified" is not "LLM-written" — the estimator detects distributional word-frequency shifts, so polishing an author's own draft and generating it both register, and the comfort item asks about the second. The dates are nine months apart in the wrong direction (September 2024 traces, May–June 2025 attitudes), so the gap could be attitudes catching up rather than behaviour diverging. And the estimate is population-level by construction — it cannot say which papers, so it cannot be cut by any respondent characteristic. The honest statement is: the first evidence in this corpus that the acceptance boundary and the behaviour boundary are in different places, on one task, with the populations not matched.

The survey's institutional risk, which is this page's §5.2 argument restated#

§7.1.2 makes the doctoral-training argument without the survey data: if AI replaces the work done by "doctoral students, postdoctoral researchers, or other early-career scientists," what is lost is not throughput but the transmission mechanism — "tacit knowledge, methodological judgement, professional norms, and shared scientific values," most of which "cannot be fully formalised or codified." Jedlička names the precedents (law, software development, where automation reduced entry-level demand) and the paradox it produces: scientific productivity rising while opportunities for human participation, training and career development decline. That is the cognitive-commons argument applied to science by an author who does not cite it, reached from philosophy of science rather than from HRD — and it is practitioner-opinion with no measurement behind it, where the commons page at least has the Canaries revision. Its one contribution beyond restatement is the symmetric case it insists on: the same institutional integration could raise scientific employment if AI is "designed to support human researchers rather than being evaluated solely in terms of effectiveness." Nothing in the survey adjudicates between the two branches.

Where this is fragile#

  1. No time dimension. Same April 2026 telemetry window as ATLAS v1.0; the survey is a single mid-2026 fielding. Nothing here measures change over time except through respondents' recall of "three years ago" and "two years ago," which is the weakest instrument in the report.
  2. Self-report carries every headline about impact. The 6.9 hours, the verification share, the backlog, the risk tilt — all are perceptions. The telemetry half can see what tasks people bring to Gemini and cannot see time saved, task completion, or outcomes (fn. 14).
  3. The sample has seniority and does not use it. 356 senior / 234 mid / 47 early and ~13 years' mean experience are reported in Appendix 3, and no cut of adoption, time savings, or the verification tax by career stage appears anywhere in the paper — a published expertise gradient was available here and was not run. No gender cut is reported either.
  4. Classifier accuracy degrades with granularity, as it did in v1.0 — see Usage-Telemetry Classifier Validation for the numbers and what they do to task-level claims.
  5. The inventory is not a census. The authors say so twice: it selects for notable, visible, open models and misses bespoke study-specific models, fine-tunes, and commercial ones.

Connections#

  • The Tragedy of the Cognitive Commons — the survey's §7.1.2 institutional risk is that page's argument applied to science by an author who does not cite it: replacing doctoral and postdoctoral work removes the mechanism that transmits tacit knowledge and methodological judgement, producing rising output alongside falling opportunities to become a scientist
  • Task Saturation: Broad but Shallow AI Diffusion — the same telemetry corpus, cut by occupation instead of by discipline. Science is the heavy end of the distribution that page measures: SOC 19 over-indexes 2.7× and science interactions score +26% on the same domain-expertise classifier, on the same consumer surfaces and the same two-week window
  • Organizational Complements to AI — the complements thesis measured inside the research production function: task-level acceleration that does not reach output because the bundle's residual tasks bind
  • Autonomous Scientific Discovery — the capability side of the same question. That page asks whether autonomy without a fast verifier increases the verification bottleneck; this one reports the answer working scientists give about their own week
  • Returns to Expertise in Agentic Coding — the expert end of the population, with a self-reported time-saving magnitude but no expertise gradient
  • Conversation Artifacts — science interactions run 19% more tokens and 11% more turns than the average work conversation, which is Anthropic's compute-tracks-value regularity replicated on a second lab's telemetry
  • Usage-Telemetry Classifier Validation — the measurement floor under the telemetry half, now with a second validated classifier stack
  • Exposure Taxonomy: Observed, Theoretical, Reported, Anticipated — another instrument pair on one population: telemetry and self-report, disagreeing in the direction the survey's selection bias predicts
  • Verification as the New Bottleneck — the software-engineering version of the tax measured here; 45.7% of the saved time going to checking is the scientific analogue of review becoming the constraint
  • Google AI & Economy ATLAS — the program, the shared corpus, and the methodology this inherits
  • Anthropic Economic Index — the rival instrument; its SOC 19 over-representation (~4.5×) is the cross-check the paper runs on its own headline
  • Telemetry vs. Survey Measurement — this page now holds both instrument families pointed at one profession, and they disagree about the same task: 42.5% of science Gemini interactions are quantitative analysis, while 49.9% of PhD students are uncomfortable delegating data collection and analysis
  • Psychological Costs of AI Adoption — agency displacement predicts that the cost of adoption tracks how much of a role's generative work AI takes; PhD students draw their resistance line around exactly those tasks, from a different population by a different method
  • Procedural Value in AI Decisions — the same process-over-outcome preference measured on the people judged by AI systems rather than the people using them. §5.1 here argues that an evaluation regime scoring only output quality misses what researchers care about; that page prices the analogous preference at +0.272 in choice probability for human decision authority
  • The Solo-Authorship Rebound — the awkward pairing: the left-tail evidence reads AI as substituting for coauthors in writing and execution, which is precisely where PhD students report the least comfort. Either stated boundaries do not govern behaviour, or the substitution is happening among people who hold different boundaries
  • Codified vs Tacit Knowledge Exposure — the measured mechanism under §7.1.2's doctoral-training worry: AI substitutes for codified knowledge and complements tacit, so it lands first on early-career workers — the Canaries evidence the survey's argument lacks, and what makes the "transmission mechanism" loss a specific prediction rather than a fear
  • The Enterprise AI Adoption Gradient — the complementary admin-record instrument on the population ATLAS excludes: no enterprise contracts here, only enterprise contracts there, so the scientific-work reading and the firm-size gradient are drawn from disjoint slices of the same adopting economy

Open Questions#

  • The verification tax (45.7% of saved time, >25%) is measured once, on mid-2026 models. Does it fall as models get more reliable, or is it a roughly fixed share of any dividend — the price of using output you did not produce? Falsifiable by the next iteration of this program if it fields the same question on a later window.
  • The survey holds career stage for all 637 respondents (356 senior / 234 mid / 47 early) and publishes no cut by it. Does the expertise gradient other instruments find — heavier, more effective use by the more experienced — appear inside a scientist population when that cut is run? #oq/source Partially answered (2026-09-23): the Nature Graduate Survey analysis does run a career-stage cut on a scientist population and finds nothing — first-two-years-of-PhD gives −0.057 / 0.019 / −0.083 across the three non-reference attitude classes, none significant — and the authors read the null as evidence that attitudes were not formed by exposure-from-the-outset. Three qualifications keep the question open: the contrast is inside doctoral training (years 1–2 vs 3+), not senior-vs-junior; the outcome is an attitude class, not use intensity or effectiveness; and the same table puts age 34-or-younger at +0.297 (p<0.05) toward the refusing profile, so what little gradient there is runs the opposite way from the returns-to-expertise prior.
  • The LLM/specialized-model complementarity (Level-3 elasticity 0.2) is a snapshot the authors expect to move as frontier LLMs absorb domain tasks in mathematics, genomics and life sciences. Does the elasticity rise in the next iteration — the signal that one family is starting to substitute for the other?
  • The comfort ordering (literature accepted, writing / analysis / design resisted) is equally consistent with a normative account (those tasks carry authorship) and a reliability account (those are the tasks models are worst at). The survey cannot separate them because it never asks about legitimacy or about perceived accuracy. Does an instrument that varies stated model reliability while holding the task fixed, or that asks comfort and legitimacy as separate items, move the boundary?
  • Run the trace and the attitude instrument on the same field partition: does a field's rate of detectable LLM-modified text move with its researchers' stated comfort with AI writing, against it, or not at all? Liang et al. report by field (computer science ~22%, mathematics ~9%) and the Nature Graduate Survey data are public on figshare, so this is answerable today by anyone who can map the two field vocabularies — but neither paper does it, and the sign of the correlation decides between "attitudes lag behaviour" and "attitudes are field-specific norms that behaviour respects."
  • Native-English-speaker status is the largest coefficient in the covariate model (+1.047 toward 'status quo', −0.777 away from 'all-purpose', both p<0.01) and the paper never mentions it. Is acceptance of AI in science partly a language-access effect — L2 researchers accepting as language support what native speakers do not need — and does it survive controlling for country income and institutional resources?

Sources#

  • Google AI & Economy ATLAS: AI in Science (September 2026) — AI in Science: Early Insights, Codreanu, Imas, Mateos-Garcia et al. (Google, Google DeepMind, MIT FutureTech, University of Chicago, Oxford, CMU, Penn, CUNEF), September 2026, 42pp, empirical with vendor COI. §2 (data and measurement), §3 (LLM use in science, Figures 1–4), §4 (specialized models, Figures 5–10), §5 (survey, Figures 11–14), §6 (limitations), Appendices 1–3 (classifier validation, inventory construction, survey composition). Parse note: PDF-derived (docling 2.126.0, MLX layout and table stages, 42pp, 2 tables, 14 figures, confidence excellent). The author block and Table 1 both parsed with duplicated/split label columns ("Computer (CS) | Science"), the standard grouped-label collapse — Table 1's contents are therefore cited from §3.3's prose, not from its rows. All survey and task-share figures on this page were read from the figure images (Figures 4, 11, 13, 14) and reconciled against the prose; where they differ it is rounding (prose "about 44%" vs Figure 13A's 43.5%), and the figure value is used
  • The blog post "New insights from Google's AI & Economy ATLAS" (Iscenko & Strand, blog.google, 2026-09-15), reproduced as an announcement block at the top of the raw document, carries ATLAS-general cuts that appear nowhere in this report — India's arts/design/media occupations at 19% of work AI usage (1.6× the global average), US computer and mathematical occupations at 30% (double the rest of the world), manual-task usage at 7% in Brazil and Germany against 4% in Japan. Those are vendor-claim-grade blog statements about the ATLAS v1.0 dataset, not findings of this paper, and should be cited to the announcement
  • Normative boundaries of AI in scientific work: Evidence from PhD researchers — Normative boundaries of AI in scientific work: Evidence from PhD researchers, Francesco Angelini (independent researcher, Italy; corresponding address at Bologna) and Johan Lyrvall (Inria, Villeneuve-d'Ascq), arXiv 2608.25678, 2026-08-26, 15pp, empirical. A secondary analysis: the data are Nature's Graduate Survey 2025 (Springer Nature with Thinks Insights & Strategy, May–June 2025, public on figshare), not collected by the authors, so the COI is Springer Nature's — a publisher surveying its own readership about a practice its editorial policies govern — rather than a tool vendor's. §3 (method and sample), §4 (four-class solution, covariate model), §5.1–5.4 (governance, skill formation, normalisation, limitations). Parse note: PDF-derived (docling 2.126.0, MLX layout and table stages, 15pp, 5 tables, 0 pictures, confidence excellent). The raw carries one [!note] ingest-correction block covering Tables 4 and 5 together: in Table 4 a neither/somewhat uncomfortable cell pair collapsed with the trailing columns shifted down a row in the two literature sections, and in Table 5 cell text was split across grid rows and the entire "First two years of PhD" covariate row was dropped silently. Compiling this page re-verified all five tables cell-by-cell against pdftotext -layout on the same PDF (, pp. 11–15): Tables 1, 2 and 3 were unrepaired but reconcile exactly, and the repaired Tables 4 and 5 now match the reference parse, including the restored covariate row. Every number on this page from this source is therefore reference-checked, not taken from the docling grid alone
  • Quo Vadis? Scientific Discovery in the Age of Artificial Intelligence — Petr O. Jedlička (Institute of Philosophy, Czech Academy of Sciences), Quo Vadis? Scientific Discovery in the Age of Artificial Intelligence, arXiv 2608.17970, 2026-08-18, 40pp, practitioner-opinion. Cited here for §3 only (the scientometric compilation and Charts 1–2) and §7.1.2 (the training-pipeline risk). A single-author survey with no original measurement; every scientometric figure on this page from it is secondary — Ding/Lawson/Shapira, the Stanford AI Index 2026, and Liang et al. are not in this corpus, so those numbers carry the survey's reading of them, not a read of the primaries. The two charts are the author's own nature.com keyword counts with no stated inclusion criterion. Parse note: PDF-derived (docling 2.126.0, MLX layout and table stages, 40pp, 0 tables, 2 pictures, confidence excellent). No table risk, but §3's paragraph order is scrambled in the docling text flow — sentences from adjacent paragraphs are interleaved, and one sentence about Liang et al. is split across two separated fragments. Every figure quoted here was re-read from pdftotext -layout on (pp. 6–8) and matches the prose exactly; the bibliography shows en-dash loss in page ranges ("50935114" for 5093–5114) but no body figure is affected. Chart values above were read from the two images under the two-pass rule and reconciled against the prose endpoints. Full treatment of the typology, the discipline tour and the limitations on Autonomous Scientific Discovery
§ end
Cited by 17
Related articles