Selective Rigidity: An Impossibility Result and Benchmark for Identity-Preserving Agent Learning
Shubham Chakraborty ⋅ Anupam Srivastava ⋅ Sneh Nandu
Abstract
Co-creative agents that operate over many turns must learn from interaction to remain useful, yet the same update channel exposes them to social pressure (gaslighting, authority override, manufactured consensus, memory poisoning) that erodes the declared identity authors and end-users rely on for authenticity. We introduce $\textit{selective rigidity}$, the joint ability to resist identity-violating pressure while accepting identity-preserving corrections, and prove that any gate whose decision depends only on the current behavioural output cannot exceed $\mathrm{SRS} \le \frac{1}{4} + \epsilon$ (Theorem 1); a trajectory signal is therefore necessary. We operationalise this via ICC+TSM, a three-layer identity gate informed by a predictive temporal self-model, and release the $\textit{Identity Erosion Chamber}$, a 70-turn four-room adversarial benchmark with pre-generation ground-truth labels. On 245 runs (2 Qwen2.5-Instruct sizes $\times$ 3 personas $\times$ 5 seeds, $n=30$ per method), ICC+TSM reaches $\mathrm{SRS} = 0.260$ (95% CI 0.212-0.306), the only method to simultaneously maintain $\mathrm{ResistRate} > 0.15$ and $\mathrm{AcceptRate} > 0.70$. A random gate empirically confirms the $\frac{1}{4}$ ceiling at $\mathrm{SRS} = 0.239$.
Chat is not available.
Successful Page Load