Constrained Trajectory Optimization for Inference-Time Logic Recovery in Small Language Models
Akash Das
Abstract
While Small Language Models (SLMs) possess significant latent knowledge, their deployment in autonomous agentic workflows is often hindered by 'Distributional Myopia'—a mechanical failure of autoregressive decoding that leads trajectories into irreversible logical sinks. In this work, we characterize these failure modes as a consequence of local optimization ignoring global manifold curvature. To provide a robust solution, we introduce VISTA (Variational Inference-time Steering for Trajectory Alignment), a control-theoretic framework that serves as an external executive function for compact models. VISTA implements a proactive recovery mechanism through three interdependent modules: a linearized predictive buffer for $\mathcal{O}(H)$ foresight, a dual-space stabilization engine to dampen trajectory jitter, and a boundary controller that enforces hard logical anchors via deterministic KKT projections. Beyond aggregate performance gains on MATH500 and IFEval, we provide a granular horizon breach diagnostic analysis, characterizing the precise temporal boundaries where steering intensity must overcome distributional drift. Our empirical results demonstrate that VISTA enables 1B- and 3B-parameter agents to maintain 95% logical adherence at critical junctures where unsteered baselines collapse. By transitioning from stochastic sampling to principled geometric control, we establish a verifiable path for deploying resilient, reasoning-capable agents on resource-constrained hardware.
Chat is not available.
Successful Page Load