Goal-Optimal Agents Necessarily Learn Predictive World Models in POMDPs with episodic resets
Keshav Goyal ⋅ suraj yadav
Abstract
We extend the recent "General agents contain world models" result of Richens et al.\ (2025) from fully observable MDPs to partially observable settings under the assumption of episodic resets. Using resets to regenerate histories and enable repeated probing, we show that any agent optimal on a sufficiently rich class of goals must implicitly encode a predictive model of its environment's observation dynamics. This extraction is robust to imperfect resets: accepting approximate neighborhoods of a target belief state rather than regenerating histories exactly introduces only a small, bounded error while yielding an exponential reduction in sample complexity.
Chat is not available.
Successful Page Load