Reasoning as a Dynamical System: Online Monitoring of LLMs with Particle Filters
Mohammed Abbas
Abstract
Online monitoring of LLM reasoning trajectories is increasingly important for inference efficiency, but existing approaches treat steps independently and provide uncalibrated confidence estimates. We model reasoning as a Switching Linear Dynamical System (SLDS) with three latent modes (Normal, Insight, Backtrack) and track it online with a Rao-Blackwellized Particle Filter (RBPF) that marginalises the continuous state analytically and requires no ground-truth labels at test time. Using per-step scores from a Process Reward Model (PRM) as observations, RBPF reaches an AUC of $0.850$ on GSM8K, comparable to strong online and offline baselines. Its main advantage is calibration: RBPF attains the lowest calibration error among all methods (ECE $0.111$), around three times lower than the strongest ranking baselines (MeanPRM $0.305$, EMA $0.311$), making its confidence estimates more directly usable downstream. For early stopping, RBPF commits a correctness verdict 12% of steps early on average, without reducing verdict accuracy. On the harder MATH, the pattern holds: RBPF is competitive on discrimination (AUPR $0.607$, the highest of all methods) and far better calibrated than the ranking baselines.
Chat is not available.
Successful Page Load