RSI for Science: A Verifier-First Framework for AI Scientists
Abstract
AI scientists are increasingly framed as recursive self-improvers: systems that generate hypotheses, choose experiments, revise tools, store lessons, and improve future campaigns. We argue that this language is misleading for empirical science unless recursion is externally verifier-governed. In empirical domains, the verifier is not a compiler, theorem checker, or game rule engine; it is a noisy, delayed, costly, and sometimes destructive physical, statistical, or procedural test. We propose Verifier-Governed Recursive Scientific Refinement (VGRSR): a standard that credits scientific recursion only when reusable objects of the scientific search process—hypotheses, tools, verifiers, memory, and campaign policy—change through cost-accounted cycles that pass independent gates. VGRSR adds four requirements to AI-scientist evaluation: external disconfirmation gates, per-cycle provenance, path-sensitive stability metrics, and verifier-efficiency accounting. We support the position with mechanism demonstrations of proxy drift and latency starvation, a five-object taxonomy, a control/risk framing using sensitivity, maximum drawdown, and drawdown CVaR, and vignettes spanning materials, cosmology, nanobiomaterials, and neurodegeneration. The practical standard is simple: no external gate, no cost-accounted cycle log, no stability report, no scientific RSI claim.