What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bu, Wendong, Wu, Yang, Yu, Qifan, Gao, Minghe, Miao, Bingchen, Zhang, Zhenkui, Pan, Kaihang, Li, Yunfei, Li, Mengze, Ji, Wei, Li, Juncheng, Tang, Siliang, Zhuang, Yueting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913888230965248
author Bu, Wendong
Wu, Yang
Yu, Qifan
Gao, Minghe
Miao, Bingchen
Zhang, Zhenkui
Pan, Kaihang
Li, Yunfei
Li, Mengze
Ji, Wei
Li, Juncheng
Tang, Siliang
Zhuang, Yueting
author_facet Bu, Wendong
Wu, Yang
Yu, Qifan
Gao, Minghe
Miao, Bingchen
Zhang, Zhenkui
Pan, Kaihang
Li, Yunfei
Li, Mengze
Ji, Wei
Li, Juncheng
Tang, Siliang
Zhuang, Yueting
contents As multimodal large language models (MLLMs) advance, MLLM-based virtual agents have demonstrated remarkable performance. However, existing benchmarks face significant limitations, including uncontrollable task complexity, extensive manual annotation with limited scenarios, and a lack of multidimensional evaluation. In response to these challenges, we introduce OmniBench, a self-generating, cross-platform, graph-based benchmark with an automated pipeline for synthesizing tasks of controllable complexity through subtask composition. To evaluate the diverse capabilities of virtual agents on the graph, we further present OmniEval, a multidimensional evaluation framework that includes subtask-level evaluation, graph-based metrics, and comprehensive tests across 10 capabilities. Our synthesized dataset contains 36k graph-structured tasks across 20 scenarios, achieving a 91\% human acceptance rate. Training on our graph-structured data shows that it can more efficiently guide agents compared to manually annotated data. We conduct multidimensional evaluations for various open-source and closed-source models, revealing their performance across various capabilities and paving the way for future advancements. Our project is available at https://omni-bench.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08933
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
Bu, Wendong
Wu, Yang
Yu, Qifan
Gao, Minghe
Miao, Bingchen
Zhang, Zhenkui
Pan, Kaihang
Li, Yunfei
Li, Mengze
Ji, Wei
Li, Juncheng
Tang, Siliang
Zhuang, Yueting
Computer Vision and Pattern Recognition
As multimodal large language models (MLLMs) advance, MLLM-based virtual agents have demonstrated remarkable performance. However, existing benchmarks face significant limitations, including uncontrollable task complexity, extensive manual annotation with limited scenarios, and a lack of multidimensional evaluation. In response to these challenges, we introduce OmniBench, a self-generating, cross-platform, graph-based benchmark with an automated pipeline for synthesizing tasks of controllable complexity through subtask composition. To evaluate the diverse capabilities of virtual agents on the graph, we further present OmniEval, a multidimensional evaluation framework that includes subtask-level evaluation, graph-based metrics, and comprehensive tests across 10 capabilities. Our synthesized dataset contains 36k graph-structured tasks across 20 scenarios, achieving a 91\% human acceptance rate. Training on our graph-structured data shows that it can more efficiently guide agents compared to manually annotated data. We conduct multidimensional evaluations for various open-source and closed-source models, revealing their performance across various capabilities and paving the way for future advancements. Our project is available at https://omni-bench.github.io/.
title What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.08933