IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Luo, Haohao, Li, Zexi, Xie, Yuexiang, Zhang, Wenhao, Li, Yaliang, Shen, Ying
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912872205910016
author Luo, Haohao
Li, Zexi
Xie, Yuexiang
Zhang, Wenhao
Li, Yaliang
Shen, Ying
author_facet Luo, Haohao
Li, Zexi
Xie, Yuexiang
Zhang, Wenhao
Li, Yaliang
Shen, Ying
contents Deep Research (DR) agents extend Large Language Models (LLMs) beyond parametric knowledge by autonomously retrieving and synthesizing evidence from large web corpora into long-form reports, enabling a long-horizon agentic paradigm. However, unlike real-time conversational assistants, DR is computationally expensive and time-consuming, creating an autonomy-interaction dilemma: high autonomy on ambiguous user queries often leads to prolonged execution with unsatisfactory outcomes. To address this, we propose IntentRL, a framework that trains proactive agents to clarify latent user intents before starting long-horizon research. To overcome the scarcity of open-ended research data, we introduce a scalable pipeline that expands a few seed samples into high-quality dialogue turns via a shallow-to-deep intent refinement graph. We further adopt a two-stage reinforcement learning (RL) strategy: Stage I applies RL on offline dialogues to efficiently learn general user-interaction behavior, while Stage II uses the trained agent and a user simulator for online rollouts to strengthen adaptation to diverse user feedback. Extensive experiments show that IntentRL significantly improves both intent hit rate and downstream task performance, outperforming the built-in clarify modules of closed-source DR agents and proactive LLM baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03468
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement Learning
Luo, Haohao
Li, Zexi
Xie, Yuexiang
Zhang, Wenhao
Li, Yaliang
Shen, Ying
Artificial Intelligence
Machine Learning
Deep Research (DR) agents extend Large Language Models (LLMs) beyond parametric knowledge by autonomously retrieving and synthesizing evidence from large web corpora into long-form reports, enabling a long-horizon agentic paradigm. However, unlike real-time conversational assistants, DR is computationally expensive and time-consuming, creating an autonomy-interaction dilemma: high autonomy on ambiguous user queries often leads to prolonged execution with unsatisfactory outcomes. To address this, we propose IntentRL, a framework that trains proactive agents to clarify latent user intents before starting long-horizon research. To overcome the scarcity of open-ended research data, we introduce a scalable pipeline that expands a few seed samples into high-quality dialogue turns via a shallow-to-deep intent refinement graph. We further adopt a two-stage reinforcement learning (RL) strategy: Stage I applies RL on offline dialogues to efficiently learn general user-interaction behavior, while Stage II uses the trained agent and a user simulator for online rollouts to strengthen adaptation to diverse user feedback. Extensive experiments show that IntentRL significantly improves both intent hit rate and downstream task performance, outperforming the built-in clarify modules of closed-source DR agents and proactive LLM baselines.
title IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement Learning
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2602.03468