Probabilistic Modeling of Latent Agentic Substructures in Deep Neural Networks
Abstract
We develop a theory of intelligent agency grounded in probabilistic modeling for neural models. Agents are represented as outcome distributions with epistemic utility given by log score, and compositions are defined through weighted logarithmic pooling that strictly improves every member's welfare. We prove that strict unanimity is impossible under linear pooling or in binary outcome spaces, but possible with three or more outcomes. Our framework admits recursive structure via cloning invariance, continuity, and openness, while tilt-based analysis rules out trivial duplication. Finally, we formalize an agentic alignment phenomenon in LLMs using our theory: eliciting a benevolent persona ("Luigi'') induces an antagonistic counterpart ("Waluigi''), while a manifest-then-suppress Waluigi strategy yields strictly larger first-order misalignment reduction than pure Luigi reinforcement alone. These results clarify how developing a principled mathematical framework for how subagents can coalesce into coherent higher-level entities provides novel implications for alignment in agentic AI systems.
Lay Summary
This paper gives a mathematical theory for thinking about AI models as made up of multiple hidden “subagents” or behavioral tendencies. Each subagent is represented as a probability distribution over possible outcomes, and the full model is represented as a special kind of combination of these distributions. The paper proves when such combinations can make all subagents better off, when they cannot, and how this depends on the number of possible outcomes and the type of averaging used. It then applies the theory to the “Waluigi effect,” arguing that strengthening an aligned persona may force the model to also reveal or strengthen an opposing anti-aligned direction, and that identifying and suppressing this direction can theoretically be more effective than simply reinforcing the good behavior.