UserBench: An Interactive Gym Environment for User-Centric Agents

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Qian, Cheng, Liu, Zuxin, Prabhakar, Akshara, Liu, Zhiwei, Zhang, Jianguo, Chen, Haolin, Ji, Heng, Yao, Weiran, Heinecke, Shelby, Savarese, Silvio, Xiong, Caiming, Wang, Huan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918107264581632
author Qian, Cheng
Liu, Zuxin
Prabhakar, Akshara
Liu, Zhiwei
Zhang, Jianguo
Chen, Haolin
Ji, Heng
Yao, Weiran
Heinecke, Shelby
Savarese, Silvio
Xiong, Caiming
Wang, Huan
author_facet Qian, Cheng
Liu, Zuxin
Prabhakar, Akshara
Liu, Zhiwei
Zhang, Jianguo
Chen, Haolin
Ji, Heng
Yao, Weiran
Heinecke, Shelby
Savarese, Silvio
Xiong, Caiming
Wang, Huan
contents Large Language Models (LLMs)-based agents have made impressive progress in reasoning and tool use, enabling them to solve complex tasks. However, their ability to proactively collaborate with users, especially when goals are vague, evolving, or indirectly expressed, remains underexplored. To address this gap, we introduce UserBench, a user-centric benchmark designed to evaluate agents in multi-turn, preference-driven interactions. UserBench features simulated users who start with underspecified goals and reveal preferences incrementally, requiring agents to proactively clarify intent and make grounded decisions with tools. Our evaluation of leading open- and closed-source LLMs reveals a significant disconnect between task completion and user alignment. For instance, models provide answers that fully align with all user intents only 20% of the time on average, and even the most advanced models uncover fewer than 30% of all user preferences through active interaction. These results highlight the challenges of building agents that are not just capable task executors, but true collaborative partners. UserBench offers an interactive environment to measure and advance this critical capability.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22034
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UserBench: An Interactive Gym Environment for User-Centric Agents
Qian, Cheng
Liu, Zuxin
Prabhakar, Akshara
Liu, Zhiwei
Zhang, Jianguo
Chen, Haolin
Ji, Heng
Yao, Weiran
Heinecke, Shelby
Savarese, Silvio
Xiong, Caiming
Wang, Huan
Artificial Intelligence
Computation and Language
Machine Learning
Large Language Models (LLMs)-based agents have made impressive progress in reasoning and tool use, enabling them to solve complex tasks. However, their ability to proactively collaborate with users, especially when goals are vague, evolving, or indirectly expressed, remains underexplored. To address this gap, we introduce UserBench, a user-centric benchmark designed to evaluate agents in multi-turn, preference-driven interactions. UserBench features simulated users who start with underspecified goals and reveal preferences incrementally, requiring agents to proactively clarify intent and make grounded decisions with tools. Our evaluation of leading open- and closed-source LLMs reveals a significant disconnect between task completion and user alignment. For instance, models provide answers that fully align with all user intents only 20% of the time on average, and even the most advanced models uncover fewer than 30% of all user preferences through active interaction. These results highlight the challenges of building agents that are not just capable task executors, but true collaborative partners. UserBench offers an interactive environment to measure and advance this critical capability.
title UserBench: An Interactive Gym Environment for User-Centric Agents
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2507.22034