RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies
Abstract
Lay Summary
Robots often struggle with tasks that require memory, such as remembering where an object was placed, how many times something happened, or what action a person demonstrated earlier. We introduce RoboMME, a benchmark for evaluating whether robot learning models can remember, reason, and act over time. RoboMME includes manipulation tasks covering different forms of memory, including object locations, event counting, object references, and imitation from demonstrations. We use it to study several ways of adding memory to vision-language-action robot models, such as text summaries, stored visual information, and recurrent memory states. Our results show that memory improves robot performance, but different memory designs work best for different tasks. RoboMME provides a controlled testbed for developing robots that can handle longer and more realistic interactions.