IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement Learning
Abstract
Deep Research (DR) agents extend Large Language Models (LLMs) beyond parametric knowledge by autonomously retrieving and synthesizing evidence from large web corpora into long-form reports, enabling a long-horizon agentic paradigm. However, unlike real-time conversational assistants, DR is computationally expensive and time-consuming, creating an autonomy-interaction dilemma: high autonomy on ambiguous user queries often leads to prolonged execution with unsatisfactory outcomes. To address this, we propose IntentRL, a framework that trains proactive agents to clarify latent user intents before starting long-horizon research. To overcome the scarcity of open-ended research data, we introduce a scalable pipeline that expands a few seed samples into high-quality dialogue turns via a shallow-to-deep intent refinement graph. We further adopt a two-stage reinforcement learning (RL) strategy: Stage I applies RL on offline dialogues to efficiently learn general user-interaction behavior, while Stage II uses the trained agent and a user simulator for online rollouts to strengthen adaptation to diverse user feedback. Extensive experiments show that IntentRL significantly improves both intent hit rate and downstream task performance, outperforming the built-in clarify modules of closed-source DR agents and proactive LLM baselines.
Lay Summary
AI agents can now act as autonomous researchers, searching the web to write comprehensive reports. However, these tasks take significant time and computing costs. If a user's initial request is vague, the agent might spend a long time generating a report that completely misses what the user actually wanted, wasting valuable resources. We tackle this problem by teaching the agent to proactively ask clarifying questions before it starts researching. Because training data for open-ended research conversations is scarce, we developed a method to automatically generate thousands of dialogue examples from a small set of initial prompts. We then trained the agent in two stages: first by learning from these generated examples, and then by practicing with a simulated user to adapt to different human responses. Our approach significantly improves the AI agents' ability to uncover hidden user goals. By asking the right questions upfront, the AI produces much better, highly tailored reports and avoids the costly mistake of researching the wrong topic.