Attractor Inversion: A Geometric Account of Adversarial Manipulation in Human Decision-Making
Leo Lorence George ⋅ Anushri Iyer ⋅ Abhishek Bakshi ⋅ Pavan Kulkarni
Abstract
We show that a compact, compositional representation of human sequential decision-making exposes a structural vulnerability that is invisible in high-dimensional opaque surrogates: adversaries exploit compositionality \emph{selectively}, targeting specific components of the representation while leaving others intact. Fitting interpretable TinyRNNs ($d=4$ hidden units, selected unanimously across all 25 cross-validation folds) to behavioral data and applying phase portrait analysis, we decompose human decision dynamics into separable input-specific fixed points. This compositional decomposition reveals that adversarial RL agents do not manipulate behavior globally; they reshape \textbf{specific components} of the attractor landscape. In the bandit task, only no-reward fixed points shift significantly (arm0/$R{=}0$: $L^*=-0.24 \to +1.11$, permutation $p<0.001$); reward contexts are unaffected ($p=0.396$). In the Go/No-Go task, nogo and go components undergo qualitatively distinct transformations: consolidation and inversion ($-2.81 \to +1.32$, $p=0.013$) vs.\ attractor fragmentation ($p=0.007$). This component-selective disruption is only detectable because the TinyRNN's low-dimensional compositional representation makes individual components readable. The opaque GRU surrogate produces the same behavioral manipulation but its representations yield no account of which components were targeted. Critically, the compositional structure of an individual's baseline representation predicts their susceptibility to component-level disruption ($r=-0.60$, $p<0.001$; slope $=-0.86$ logits/logit), enabling safety monitoring via compositional drift detection before behavioral effects become observable.
Chat is not available.
Successful Page Load