ProAct: A Benchmark and Multimodal Framework for Structure-Aware Proactive Response

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhu, Xiaomeng, Zhu, Fengming, Zhou, Weijie, Tian, Ye, Hu, Zhenlin, Huang, Yufei, Guo, Yuchun, Wu, Xinyu, Zhang, Zhengyou, Lin, Fangzhen, Xiong, Xuantang
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911420510109696
author Zhu, Xiaomeng
Zhu, Fengming
Zhou, Weijie
Tian, Ye
Hu, Zhenlin
Huang, Yufei
Guo, Yuchun
Wu, Xinyu
Zhang, Zhengyou
Lin, Fangzhen
Xiong, Xuantang
author_facet Zhu, Xiaomeng
Zhu, Fengming
Zhou, Weijie
Tian, Ye
Hu, Zhenlin
Huang, Yufei
Guo, Yuchun
Wu, Xinyu
Zhang, Zhengyou
Lin, Fangzhen
Xiong, Xuantang
contents While passive agents merely follow instructions, proactive agents align with higher-level objectives, such as assistance and safety by continuously monitoring the environment to determine when and how to act. However, developing proactive agents is hindered by the lack of specialized resources. To address this, we introduce ProAct-75, a benchmark designed to train and evaluate proactive agents across diverse domains, including assistance, maintenance, and safety monitoring. Spanning 75 tasks, our dataset features 91,581 step-level annotations enriched with explicit task graphs. These graphs encode step dependencies and parallel execution possibilities, providing the structural grounding necessary for complex decision-making. Building on this benchmark, we propose ProAct-Helper, a reference baseline powered by a Multimodal Large Language Model (MLLM) that grounds decision-making in state detection, and leveraging task graphs to enable entropy-driven heuristic search for action selection, allowing agents to execute parallel threads independently rather than mirroring the human's next step. Extensive experiments demonstrate that ProAct-Helper outperforms strong closed-source models, improving trigger detection mF1 by 6.21%, saving 0.25 more steps in online one-step decision, and increasing the rate of parallel actions by 15.58%.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03430
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ProAct: A Benchmark and Multimodal Framework for Structure-Aware Proactive Response
Zhu, Xiaomeng
Zhu, Fengming
Zhou, Weijie
Tian, Ye
Hu, Zhenlin
Huang, Yufei
Guo, Yuchun
Wu, Xinyu
Zhang, Zhengyou
Lin, Fangzhen
Xiong, Xuantang
Robotics
While passive agents merely follow instructions, proactive agents align with higher-level objectives, such as assistance and safety by continuously monitoring the environment to determine when and how to act. However, developing proactive agents is hindered by the lack of specialized resources. To address this, we introduce ProAct-75, a benchmark designed to train and evaluate proactive agents across diverse domains, including assistance, maintenance, and safety monitoring. Spanning 75 tasks, our dataset features 91,581 step-level annotations enriched with explicit task graphs. These graphs encode step dependencies and parallel execution possibilities, providing the structural grounding necessary for complex decision-making. Building on this benchmark, we propose ProAct-Helper, a reference baseline powered by a Multimodal Large Language Model (MLLM) that grounds decision-making in state detection, and leveraging task graphs to enable entropy-driven heuristic search for action selection, allowing agents to execute parallel threads independently rather than mirroring the human's next step. Extensive experiments demonstrate that ProAct-Helper outperforms strong closed-source models, improving trigger detection mF1 by 6.21%, saving 0.25 more steps in online one-step decision, and increasing the rate of parallel actions by 15.58%.
title ProAct: A Benchmark and Multimodal Framework for Structure-Aware Proactive Response
topic Robotics
url https://arxiv.org/abs/2602.03430