MEME: Multi-Entity and Evolving Memory Evaluation
Seokwon Jung ⋅ Alexander Rubinstein ⋅ Arnas Uselis ⋅ Sangdoo Yun ⋅ Seong Joon Oh
Abstract
LLM-based agents increasingly operate in persistent environments where they must store, update, and reason over information across many sessions. Yet existing memory benchmarks evaluate updates only for independent entities, without testing whether a system can propagate the consequences of a changed fact through dependent entities or recognize when a previously valid answer has become uncertain. We introduce MEME, a benchmark that organizes memory evaluation along two orthogonal dimensions: entity scope and temporal dynamics. It defines six tasks per quadrant, including two novel ones, Cascade and Absence, that no prior benchmark covers. Evaluating six systems spanning three architectural paradigms on 100 controlled episodes, we find that all systems collapse on dependency reasoning under the default configuration (Cascade: 2.9\%, Absence: 1.4\% average) despite adequate static retrieval performance. Prompt optimization, reduced filler noise, and most stronger LLMs fail to close this gap. Only a file-based agent paired with Claude Opus 4.7 partially closes the gap, but at $\sim$70$\times$ the baseline cost, indicating closure currently depends on configurations that are not practical at scale. Code is available at https://anonymous.4open.science/r/MEME-0612 and the dataset at https://huggingface.co/datasets/meme-benchmark/MEME
Chat is not available.
Successful Page Load