A Minimal Decision Capacity Threshold Prevents Catastrophic Exploitation in Self-Play RL
Abstract
We show that a minimal threshold in decision capacity determines whether self-play reinforcement learning agents collapse under asymmetric rule perturbations. In Kuhn and Leduc Poker, we remove Player 0’s ability to bet or raise—either at all decision nodes or only at the opening move. Across five seeds with paired card deals: (i) complete removal causes adaptive Q-learning to collapse to near-maximal exploitation (Kuhn: −0.93; Leduc: −0.31), while a frozen Q-learning baseline stays near −0.14, confirming the collapse is co-adaptation-driven; (ii) preserving a single decision point stabilises Q-learning near Nash equilibrium (Kuhn: −0.07; Leduc: −0.10); (iii) the pattern is timing-invariant: early, mid, and late perturbation produce identical collapse severity; (iv) collapse is fast, occurring within four episodes on average in Kuhn. These results reveal a structural instability in learning dynamics: equilibrium behavior becomes unsustainable once agents lose all contingent responses. This establishes a sharp threshold effect between zero and minimal decision capacity that is game-general, not Kuhn-specific.