ProAgentBench: Evaluating LLM Agents for Proactive Assistance with Real-World Data

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tang, Yuanbo, Tang, Huaze, Cao, Tingyu, Nguyen, Lam, Zhang, Anping, Cao, Xinwen, Liu, Chunkang, Ding, Wenbo, Li, Yang
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914315944067072
author Tang, Yuanbo
Tang, Huaze
Cao, Tingyu
Nguyen, Lam
Zhang, Anping
Cao, Xinwen
Liu, Chunkang
Ding, Wenbo
Li, Yang
author_facet Tang, Yuanbo
Tang, Huaze
Cao, Tingyu
Nguyen, Lam
Zhang, Anping
Cao, Xinwen
Liu, Chunkang
Ding, Wenbo
Li, Yang
contents Proactive agents that anticipate user intentions without explicit prompts represent a significant evolution in human-AI interaction, promising to reduce cognitive load and streamline workflows. However, existing datasets suffer from two critical deficiencies: (1) reliance on LLM-synthesized data that fails to capture authentic human decision-making patterns, and (2) focus on isolated tasks rather than continuous workflows, missing the pre-assistance behavioral context essential for learning proactive intervention signals. To address these gaps, we introduce ProAgentBench, a rigorous benchmark for proactive agents in working scenarios. Our contributions include: (1) a hierarchical task framework that decomposes proactive assistance into timing prediction and assist content generation; (2) a privacy-compliant dataset with 28,000+ events from 500+ hours of real user sessions, preserving bursty interaction patterns (burstiness B=0.787) absent in synthetic data; and (3) extensive experiments that evaluates LLM- and VLM-based baselines. Numerically, we showed that long-term memory and historical context significantly enhance prediction accuracy, while real-world training data substantially outperforms synthetic alternatives. We release our dataset and code at https://anonymous.4open.science/r/ProAgentBench-6BC0.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04482
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ProAgentBench: Evaluating LLM Agents for Proactive Assistance with Real-World Data
Tang, Yuanbo
Tang, Huaze
Cao, Tingyu
Nguyen, Lam
Zhang, Anping
Cao, Xinwen
Liu, Chunkang
Ding, Wenbo
Li, Yang
Human-Computer Interaction
Proactive agents that anticipate user intentions without explicit prompts represent a significant evolution in human-AI interaction, promising to reduce cognitive load and streamline workflows. However, existing datasets suffer from two critical deficiencies: (1) reliance on LLM-synthesized data that fails to capture authentic human decision-making patterns, and (2) focus on isolated tasks rather than continuous workflows, missing the pre-assistance behavioral context essential for learning proactive intervention signals. To address these gaps, we introduce ProAgentBench, a rigorous benchmark for proactive agents in working scenarios. Our contributions include: (1) a hierarchical task framework that decomposes proactive assistance into timing prediction and assist content generation; (2) a privacy-compliant dataset with 28,000+ events from 500+ hours of real user sessions, preserving bursty interaction patterns (burstiness B=0.787) absent in synthetic data; and (3) extensive experiments that evaluates LLM- and VLM-based baselines. Numerically, we showed that long-term memory and historical context significantly enhance prediction accuracy, while real-world training data substantially outperforms synthetic alternatives. We release our dataset and code at https://anonymous.4open.science/r/ProAgentBench-6BC0.
title ProAgentBench: Evaluating LLM Agents for Proactive Assistance with Real-World Data
topic Human-Computer Interaction
url https://arxiv.org/abs/2602.04482