Bridging Game Theory and Transformer Routing: Mean Field Equilibria for Mixture of Experts
Nevroz Sen
Abstract
Current Mixture of Experts (MoE) routing relies on heuristic load-balancing objectives whose equilibrium structure is not characterized, leading to token dropping and suboptimal load distribution. We reformulate MoE routing as a mean field game where tokens compete for expert capacity, deriving a Nash equilibrium routing policy jointly optimizing quality and load balance. For linear congestion costs, we derive a closed-form best response converging in $O(\log 1/\varepsilon)$ iterations, establish existence and uniqueness, and prove $O(1/\sqrt{N}+\sqrt{N/T})$ finite-expert approximation bounds. Empirically, our Capacity-Aware MFG router improves WikiText-103 perplexity over Switch by 12.4\% (at $1.65\times$ cost) while eliminating irreversible token drops. Perplexity is largely insensitive to the congestion penalty across a $25\times$ range, with settings inside the contractive regime achieving nearly identical performance. On heterogeneous multi-domain data (Text+Code and Text+Code+Math), the MFG advantage remains consistent at 9--12\% over Switch and improves over dense random routing. With sparse-aware training, the equilibrium projects to Top-1 routing, recovering a 2.7\% gain at identical inference cost.
Chat is not available.
Successful Page Load