ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ai, Jiaxin, Feng, Yukang, Zhang, Fanrui, Sun, Jianwen, Li, Zizhen, Li, Chuanhao, Chang, Yifan, Wu, Wenxiao, Wang, Ruoxi, Zhai, Mingliang, Zhang, Kaipeng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915709568679936
author Ai, Jiaxin
Feng, Yukang
Zhang, Fanrui
Sun, Jianwen
Li, Zizhen
Li, Chuanhao
Chang, Yifan
Wu, Wenxiao
Wang, Ruoxi
Zhai, Mingliang
Zhang, Kaipeng
author_facet Ai, Jiaxin
Feng, Yukang
Zhang, Fanrui
Sun, Jianwen
Li, Zizhen
Li, Chuanhao
Chang, Yifan
Wu, Wenxiao
Wang, Ruoxi
Zhai, Mingliang
Zhang, Kaipeng
contents Multimodal agents are making rapid progress on general computer-use tasks, yet existing benchmarks remain largely confined to browsers and basic desktop applications, falling short in professional software workflows that dominate real-world scientific and industrial practice. To close this gap, we introduce ProSoftArena, a benchmark and platform specifically for evaluating multimodal agents in professional software environments. We establish the first capability hierarchy tailored to agent use of professional software and construct a benchmark of 436 realistic work and research tasks spanning 6 disciplines and 13 core professional applications. To ensure reliable and reproducible assessment, we build an executable real-computer environment with an execution-based evaluation framework and uniquely incorporate a human-in-the-loop evaluation paradigm. Extensive experiments show that even the best-performing agent attains only a 24.4\% success rate on L2 tasks and completely fails on L3 multi-software workflow. In-depth analysis further provides valuable insights for addressing current agent limitations and more effective design principles, paving the way to build more capable agents in professional software settings. This project is available at: https://prosoftarena.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2601_02399
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
Ai, Jiaxin
Feng, Yukang
Zhang, Fanrui
Sun, Jianwen
Li, Zizhen
Li, Chuanhao
Chang, Yifan
Wu, Wenxiao
Wang, Ruoxi
Zhai, Mingliang
Zhang, Kaipeng
Software Engineering
Artificial Intelligence
Multimodal agents are making rapid progress on general computer-use tasks, yet existing benchmarks remain largely confined to browsers and basic desktop applications, falling short in professional software workflows that dominate real-world scientific and industrial practice. To close this gap, we introduce ProSoftArena, a benchmark and platform specifically for evaluating multimodal agents in professional software environments. We establish the first capability hierarchy tailored to agent use of professional software and construct a benchmark of 436 realistic work and research tasks spanning 6 disciplines and 13 core professional applications. To ensure reliable and reproducible assessment, we build an executable real-computer environment with an execution-based evaluation framework and uniquely incorporate a human-in-the-loop evaluation paradigm. Extensive experiments show that even the best-performing agent attains only a 24.4\% success rate on L2 tasks and completely fails on L3 multi-software workflow. In-depth analysis further provides valuable insights for addressing current agent limitations and more effective design principles, paving the way to build more capable agents in professional software settings. This project is available at: https://prosoftarena.github.io.
title ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2601.02399