Implicit Neural Representations of Individual Behavior
Abstract
We study policy representation learning from unlabeled multi-policy behavioral data, where each episode is generated by a fixed but unobserved policy. We introduce Behavioral INR, a self-supervised generative model that adapts implicit neural representations from vision to behavior: instead of mapping coordinates to RGB values, it represents a policy as a state-action function mapping states to actions. An episode-level latent modulates this function through FiLM layers, yielding a generative prior over policies and enabling policy identity inference without labels. We also define policy-level out-of-distribution (OOD) shifts along state- and action-distribution axes, which arise when policies overlap but are not captured by standard agent- or environment-level OOD settings. Across synthetic GRF data, Seek-Avoid, MuJoCo, chess, Formula 1, and robotic manipulation, Behavioral INR most consistently improves policy identifiability in harder continuous state-action settings, especially when longer episodes, more policies, and OOD splits reduce marginal shortcuts.