Conversational Hallucination Drift: An Episodic Retrieval Framework and Error Taxonomy for Long-Term Memory Evaluation
Kamesh Rajeshkanna ⋅ P Sivadhanushya
Abstract
Reliable long-term memory is a prerequisite for deploying Large Language Models (LLMs) as effective agents in multi-session, real-world settings. Yet despite growing interest in retrieval-augmented memory for agentic systems, the failure modes underlying memory errors remain poorly characterised, limiting principled progress toward robust agents. We introduce the Conversational Hallucination Drift (CHD) taxonomy, which classifies memory failures into four mechanistically distinct types: commission, omission, distortion, and confabulation. Evaluating on LongMemEval using an open episodic retrieval framework with Qwen3-32B as the reasoning backbone, we reach 24.9% accuracy with recall and scout summaries, performing on par with a commercial GPT-3.5 memory baseline. Three independent annotators labelled all 500 evaluation questions (Fleiss' $\kappa = 0.91$), revealing a key structural dissociation: confabulation errors occur even when the correct answer is already present in the loaded context, while omission errors correspond to cases where the model honestly signals retrieval failure. This identifies two structurally independent failure modes: omission, which is addressable through improved retrieval coverage, and confabulation, which persists despite successful retrieval and therefore implicates generation faithfulness as a distinct, retrieval-resistant bottleneck. Our findings suggest that scaling retrieval alone is insufficient for trustworthy agentic memory, and that faithfulness-aware generation warrants prioritisation in future agent architectures.
Chat is not available.
Successful Page Load