Physics Is All You Need? A Case Study in Physicist-Supervised AI Development of Scientific Software
Nhat-Minh Nguyen
Abstract
Are AI agents tools, co-authors, or researchers? We present a quantified case study ($N=1$): a physicist supervising an AI coding agent (Claude Code, Sonnet and Opus models) over 12 work days and 57 sessions to build a differentiable one-loop perturbation theory (a next-to-leading-order calculation for predicting galaxy clustering) module in JAX, CLAX-PT ($\sim$2,100 lines, validated to $\lesssim$1\% accuracy against the established C reference CLASS-PT). We documented 15 supervision events during the v0.1.0 development window and classified each by intervention level (Appendix A). The agent resolved ten autonomously (convention errors, algorithm transcription, numerical coefficients) by iterating against oracle test suites. Two more were accelerated by the physicist spotting magnitude discrepancies invisible to shape-based comparisons. The three it could not---all evaded oracle detection---share a common property: the agent treated symptom reduction as equivalent to root-cause resolution. It spent 33 of the 57 sessions adjusting coefficients within a code architecture that could not represent the target physics, and could not re-evaluate its choice of CLASS-PT branch even when the physicist explicitly prompted reconsideration; only an injected physics concept (anisotropic BAO damping) triggered the redesign. Separately, the agent committed a calibrated scalar correction that passed all oracle tests, but the value corresponded to no quantity in the reference theory and would produce wrong predictions at any other cosmology. The fudge factor was caught and replaced within the same session. Three supervision practices, developed iteratively, proved critical for catching what oracle tests missed: testing at diverse parameter points beyond the fiducial calibration; shared changelogs that surfaced stalled exploration across sessions; and an explicit rule against unphysical numerical patches. In this case study, the design of these supervision protocols, not model capability, was the primary factor in whether the agent's output was trustworthy. Closing the gap we observed would require agents that can propose architectural alternatives rather than optimize within a given structure, and distinguish predictive adequacy from explanatory correctness. The agent in this case study did not exhibit these capabilities, and they are not obviously addressed by scaling alone.
Chat is not available.
Successful Page Load