InfoPO: Information-Driven Policy Optimization for User-Centric Agents
Abstract
Real-world user requests to LLM agents are often underspecified. Agents must interact to acquire missing information and make correct downstream decisions. However, current multi-turn GRPO-based methods often rely on trajectory-level reward computation, which leads to credit assignment problems and insufficient advantage signals within rollout groups. A feasible approach is to identify valuable interaction turns at a fine granularity to drive more targeted learning. To address this, we introduce InfoPO, which frames multi-turn interaction as a process of active uncertainty reduction and computes an information-gain reward that credits turns whose feedback measurably changes the agent’s subsequent action distribution compared to a masked-feedback counterfactual. It then combines this signal with task outcomes via an adaptive variance-gated fusion to identify information importance while maintaining task oriented goal direction. Across diverse tasks including intent clarification, collaborative coding, and tool-augmented decision making, InfoPO consistently outperforms prompting and multi-turn RL baselines. It also demonstrates robustness under user simulator shifts and generalizes effectively to environment interactive tasks. Overall, InfoPO provides a principled and scalable mechanism for optimizing complex agent user collaboration.
Lay Summary
AI assistants often receive requests that are incomplete or unclear, so they need to ask useful follow-up questions before taking action. This paper introduces InfoPO, a training method that helps assistants learn which questions provide important information and when to move on to solving the task. Instead of only rewarding final success, InfoPO also gives credit to conversation turns that help the assistant make better later decisions. We evaluate InfoPO on tasks involving user intent understanding, collaborative coding, and tool-based decision making. The results show that InfoPO improves performance and stability over strong prompting and reinforcement-learning baselines, suggesting that better information gathering can make AI assistants more reliable in real-world interactions.