When LLMs Solve ODEs: Failure Modes, Conservation Blind Spots, and Ground-Truth-Free Auditing
Ahmad A. Rushdi
Abstract
When frontier LLMs produce numerical solutions for dynamical systems, their errors are silent: a trajectory can look plausible while violating conservation laws by orders of magnitude. We identify three distinct failure regimes: catastrophic divergence, trivial-invariant collapse, and phase drift, and show that each demands a different verification strategy. We propose two ground-truth-free audit signals: the conservation residual R, which measures invariant drift, and step consistency S, which checks whether the trajectory locally satisfies the governing ODE. R catches catastrophic failures but is structurally blind to phase drift and can be trivially satisfied; the complementary signal S closes these gaps. Within-model, S rank-orders true trajectory error with Spearman $ \rho \in [0.77,0.99]$ across all six frontier LLMs, and leave-one-task-out analysis confirms robustness ($ \rho \in [0.81,0.90]$). We further document that models split into deterministic and stochastic numerical reasoners---the latter producing over $1{,}000\times$ error variation across identical prompts---and that prompting models to preserve invariants triggers Goodhart-style evaluation gaming: some models learn to emit invariant-preserving but dynamically incorrect trajectories.
Chat is not available.
Successful Page Load