Generalization vs. Memorization in Partial Differential Equation Emulators: Or, Training Dynamics of Cross-Time and Cross-Class Gradient Alignment
Abstract
Neural emulators of solution operators to partial differential equations capable of reliable generalization stand to accelerate scientific discovery by extrapolating beyond training regimes and performing over varied initial condition classes. Assessing model generalization capabilities requires distinguishing solution operator-consistent behavior from memorization of feature derived correlations, which can be achieved by using influence functions to characterize how information is shared across examples with different temporal indices and/or of different initial conditions. We compute the hat-matrix metric-weighted gradient alignment by inverting the relevant Hessian restricted to the mini-batch gradient-spanned subspace, enabling exact measurement of cross-time and cross-class coherence during training. This diagnostic is architecture-agnostic and can be efficiently computed. Our results indicate that gradient alignment between training examples is governed by their feature-space similarity. We find that gradient coherence across initial condition classes is suppressed due to feature-space dissimilarity. Moreover, we quantify the temporal decay of gradient alignment induced by standard one-step autoregressive training, revealing that gradient coherence is not extended in time, which limits long-horizon extrapolation performance. Cross-time influence exhibits a negative correlation with rollout error across horizons, indicating that the measured response geometry contains information predictive of inference-time degradation.