Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Miao, Bingchen, Wu, Yang, Gao, Minghe, Yu, Qifan, Bu, Wendong, Zhang, Wenqiao, Li, Yunfei, Tang, Siliang, Chua, Tat-Seng, Li, Juncheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911018002677760
author Miao, Bingchen
Wu, Yang
Gao, Minghe
Yu, Qifan
Bu, Wendong
Zhang, Wenqiao
Li, Yunfei
Tang, Siliang
Chua, Tat-Seng
Li, Juncheng
author_facet Miao, Bingchen
Wu, Yang
Gao, Minghe
Yu, Qifan
Bu, Wendong
Zhang, Wenqiao
Li, Yunfei
Tang, Siliang
Chua, Tat-Seng
Li, Juncheng
contents The development of Generalist Virtual Agents (GVAs) has shown significant promise in autonomous task execution. However, current training paradigms face critical limitations, including reliance on outcome supervision and labor-intensive human annotations. To address these challenges, we propose Similar, a Step-Wise Multi-Dimensional Generalist Reward Model, which offers fine-grained signals for agent training and can choose better action for inference-time scaling. Specifically, we begin by systematically defining five dimensions for evaluating agent actions. Building on this framework, we design an MCTS-P algorithm to automatically collect and annotate step-wise, five-dimensional agent execution data. Using this data, we train Similar with the Triple-M strategy. Furthermore, we introduce the first benchmark in the virtual agent domain for step-wise, multi-dimensional reward model training and evaluation, named SRM. This benchmark consists of two components: SRMTrain, which serves as the training set for Similar, and SRMEval, a manually selected test set for evaluating the reward model. Experimental results demonstrate that Similar, through its step-wise, multi-dimensional assessment and synergistic gain, provides GVAs with effective intermediate signals during both training and inference-time scaling. The project is available at https://github.com/antgroup/Similar.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18665
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark
Miao, Bingchen
Wu, Yang
Gao, Minghe
Yu, Qifan
Bu, Wendong
Zhang, Wenqiao
Li, Yunfei
Tang, Siliang
Chua, Tat-Seng
Li, Juncheng
Computer Vision and Pattern Recognition
The development of Generalist Virtual Agents (GVAs) has shown significant promise in autonomous task execution. However, current training paradigms face critical limitations, including reliance on outcome supervision and labor-intensive human annotations. To address these challenges, we propose Similar, a Step-Wise Multi-Dimensional Generalist Reward Model, which offers fine-grained signals for agent training and can choose better action for inference-time scaling. Specifically, we begin by systematically defining five dimensions for evaluating agent actions. Building on this framework, we design an MCTS-P algorithm to automatically collect and annotate step-wise, five-dimensional agent execution data. Using this data, we train Similar with the Triple-M strategy. Furthermore, we introduce the first benchmark in the virtual agent domain for step-wise, multi-dimensional reward model training and evaluation, named SRM. This benchmark consists of two components: SRMTrain, which serves as the training set for Similar, and SRMEval, a manually selected test set for evaluating the reward model. Experimental results demonstrate that Similar, through its step-wise, multi-dimensional assessment and synergistic gain, provides GVAs with effective intermediate signals during both training and inference-time scaling. The project is available at https://github.com/antgroup/Similar.
title Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.18665