Mitigating Over-Personalization in Language Models via Structured Memory
Abdal Hakeem Hannoon ⋅ Andrew Zhao ⋅ Mihir Narayan ⋅ Sharvin Goyal ⋅ Ivaxi Sheth
Abstract
Conversational language models now ship with persistent user memory, prepended to the system prompt as a flat textual list. Recent work has shown this format induces cross-domain leakage and memory-induced sycophancy. We investigate representation-level mitigations: reorganizing the same memory set at inference time into modular, domain-structured representations -- fixed-domain partitions, dynamic partitions, or a two-level memory tree -- without changing the underlying model or memory content. On PersistBench across seven frontier models, fixed partitioning reduces cross-domain leakage in six of seven models, while dynamic partitioning improves all seven and lowers leakage by $~8.8\%$ on average relative to the flat baseline, while preserving desired personalization. These transformations also stack on top of some prompt-based defenses. Our results indicate that imposing compositional structure on the memory block is a useful inductive bias for how models integrate persistent context, and motivate further work on structured memory representations.
Chat is not available.
Successful Page Load