Grounded autonomous scrutiny at scale: emergent critique from reproduction of published computational physics papers
Abstract
Autonomous LLM agents now produce complete research artifacts in machine-learning sandboxes with minimal human input. Follow-up evaluations have raised concerns about reliability — notably unverified numerical results and execution errors — yet these concerns often survive conventional peer review, because human reviewers do not, in practice, re-execute the underlying calculations as part of the review process. The missing capability is grounded scrutiny: an agent that reads a published paper, autonomously plans and re-executes its underlying calculations end-to-end, identifies what does and does not hold, and extends it where warranted. We establish this capability for computational physics, where re-runnable physical ground truth makes scrutiny falsifiable. At scale, across 111 open-access computational physics papers, an agent autonomously runs the read–plan–compute–compare pipeline; instructed only to reproduce, it nevertheless raises substantive methodological concerns on ~42% of papers — 97.7% of which emerge only after execution, against a reading-only ceiling of 0.9%. Critique emerges from reproduction, not from reading. In depth, on a published peer-reviewed paper on multiscale simulation of a 2D-material MOSFET, the agent runs new calculations missing from the original and produces, unsupervised, a publishable Comment — composed, figured, typeset, and PDF-iterated — that revises the paper's headline conclusion, including findings absent from the published 21-reviewer peer review.