DOI: 10.1073/pnas.2618638123 ISSN: 0027-8424
Scientific production in the era of large language models: Outcome-triggered treatment timing and spurious event-study dynamics
Thomas Renault, Antonin Bergeaud, Clément Bosquet
Large language models (LLMs) are increasingly used in scientific writing, but their effect on individual productivity is difficult to identify because adoption is rarely directly observed. [K. Kusumegi
et al.
,
Science
390
, 1240–1243 (2025)] infer adoption from the first paper detected as LLM-assisted and report large productivity gains after adoption. We show that this treatment-timing rule mechanically generates positive event-study dynamics even in the absence of any causal effect. Because high-output months are more likely to produce a detected paper, treatment assignment becomes intrinsically linked to productivity. Using reconstructed arXiv data, we show that random treatment assignments, neutral-keyword triggers, inverted treatment, and pre-ChatGPT placebo periods all generate similar dynamics. Simulations with no treatment effect also reproduce the same posttreatment patterns reported in K. Kusumegi
et al.
,
Science
390
, 1240–1243 (2025). Our results demonstrate that first-detection timing alone creates spurious evidence of productivity gains from LLM adoption.