Mem-T: Densifying Rewards for Long-Horizon Memory Agents
Abstract
Lay Summary
It is crucial to construct a memory system that encompasses a wide variety of functions and to train it on diverse memory operations. However, the processes of memory construction and retrieval are excessively long-range and sparse, making it difficult for feedback from a single question to reflect the quality of relevant memory construction. To address this, we propose Mem-T, an autonomous memory agent that interfaces with a lightweight hierarchical memory database. Furthermore, we introduce MoT-GRPO, which constructs dense rewards through a Memory Operation Tree (MoT) and identifies critical retrieval operations. By using memory items and evidence as pivots, this approach propagates reasoning feedback to evaluate the quality of the corresponding memory construction operations. Consequently, this achieves optimization for both memory construction and memory retrieval operations within the hierarchical memory systems.