Deterministic State-Space Governance of LLMs via Multi-Objective Optimization Envelopes and Jacobian Steering
Akash Das ⋅ Ishan Roy
Abstract
As foundation models scale and integrate into complex workflows, securing them against sophisticated adversarial attacks and structural jailbreaks remains a critical challenge. Traditional categorical guardrails are fundamentally insufficient, frequently collapsing under adversarial strain or enforcing a severe alignment tax that degrades the model's native reasoning capabilities. To address this, we introduce the Multi-Objective Optimization Envelope (MOE), a continuous geometric governance framework that transitions AI safety from reactive post-hoc filtering to proactive state-space steering. Utilizing Data Envelopment Analysis (DEA), we construct a convex relaxation of a foundation model's optimal reasoning-safety frontier. Our empirical evaluations demonstrate that adversarial drift—such as stealth jailbreaks and information hazards—induces a measurable structural conflict, quantifiable as a Mahalanobis-scaled Efficiency Deficit (mean $\beta=3.36$). To neutralize these vulnerabilities in real-time, MOE applies Projected Subgradient Steering to the latent logit distribution, deterministically correcting the generative trajectory. This approach successfully secures foundation models against adversarial exploits while mitigating the severe latency and utility degradation associated with legacy safety patches.
Chat is not available.
Successful Page Load