ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Ai, Jiaxin, Feng, Yukang, Zhang, Fanrui, Sun, Jianwen, Li, Zizhen, Li, Chuanhao, Chang, Yifan, Wu, Wenxiao, Wang, Ruoxi, Zhai, Mingliang, Zhang, Kaipeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
by: Sun, Jianwen, et al.
Published: (2025)
by: Sun, Jianwen, et al.
Published: (2025)
Closing the Expression Gap in LLM Instructions via Socratic Questioning
by: Sun, Jianwen, et al.
Published: (2025)
by: Sun, Jianwen, et al.
Published: (2025)
From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration
by: Sun, Jianwen, et al.
Published: (2025)
by: Sun, Jianwen, et al.
Published: (2025)
IA-T2I: Internet-Augmented Text-to-Image Generation
by: Li, Chuanhao, et al.
Published: (2025)
by: Li, Chuanhao, et al.
Published: (2025)
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
by: Feng, Yukang, et al.
Published: (2025)
by: Feng, Yukang, et al.
Published: (2025)
MeepleLM: A Virtual Playtester Simulating Diverse Subjective Experiences
by: Li, Zizhen, et al.
Published: (2026)
by: Li, Zizhen, et al.
Published: (2026)
World Craft: Agentic Framework to Create Visualizable Worlds via Text
by: Sun, Jianwen, et al.
Published: (2026)
by: Sun, Jianwen, et al.
Published: (2026)
InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles
by: Li, Zizhen, et al.
Published: (2025)
by: Li, Zizhen, et al.
Published: (2025)
ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges
by: Ai, Jiaxin, et al.
Published: (2025)
by: Ai, Jiaxin, et al.
Published: (2025)
AutoBG: A Board Game Design Assistant with Interactive Ideation, Iterative Rulebook Generation, and Individualized Feedback
by: Li, Zizhen, et al.
Published: (2026)
by: Li, Zizhen, et al.
Published: (2026)
Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection
by: Zhang, Fanrui, et al.
Published: (2025)
by: Zhang, Fanrui, et al.
Published: (2025)
SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model
by: Chang, Yifan, et al.
Published: (2025)
by: Chang, Yifan, et al.
Published: (2025)
LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces
by: Feng, Yukang, et al.
Published: (2026)
by: Feng, Yukang, et al.
Published: (2026)
MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models
by: Zhou, Pengfei, et al.
Published: (2025)
by: Zhou, Pengfei, et al.
Published: (2025)
MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams
by: Zhou, Pengfei, et al.
Published: (2025)
by: Zhou, Pengfei, et al.
Published: (2025)
Sekai: A Video Dataset towards World Exploration
by: Li, Zhen, et al.
Published: (2025)
by: Li, Zhen, et al.
Published: (2025)
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision-Language Models
by: Liu, Shuo, et al.
Published: (2024)
by: Liu, Shuo, et al.
Published: (2024)
LaGen: Towards Autoregressive LiDAR Scene Generation
by: Zhou, Sizhuo, et al.
Published: (2025)
by: Zhou, Sizhuo, et al.
Published: (2025)
Hierarchically Reconfigurable Soft Robots with Reprogrammable Multimodal Actuation
by: Fuyi Fang, et al.
Published: (2024)
by: Fuyi Fang, et al.
Published: (2024)
Lactate‐induced metabolic reprogramming of TAMs impairs antigen presentation capacity via C/EBPα–CD74 axis in oral squamous cell carcinoma
by: Mengyao Wang, et al.
Published: (2026)
by: Mengyao Wang, et al.
Published: (2026)
MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification
by: Xu, Zhaopan, et al.
Published: (2025)
by: Xu, Zhaopan, et al.
Published: (2025)
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
by: Mao, Xiaofeng, et al.
Published: (2026)
by: Mao, Xiaofeng, et al.
Published: (2026)
SVBench: Evaluation of Video Generation Models on Social Reasoning
by: Peng, Wenshuo, et al.
Published: (2025)
by: Peng, Wenshuo, et al.
Published: (2025)
Hierarchical Information Enhancement Network for Cascade Prediction in Social Networks
by: Zhang, Fanrui, et al.
Published: (2024)
by: Zhang, Fanrui, et al.
Published: (2024)
TP-Eval: Tap Multimodal LLMs' Potential in Evaluation by Customizing Prompts
by: Xie, Yuxuan, et al.
Published: (2024)
by: Xie, Yuxuan, et al.
Published: (2024)
LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition
by: Zheng, Yushuo, et al.
Published: (2025)
by: Zheng, Yushuo, et al.
Published: (2025)
WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG
by: Li, Zhen, et al.
Published: (2026)
by: Li, Zhen, et al.
Published: (2026)
PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models
by: Xu, Zhaopan, et al.
Published: (2025)
by: Xu, Zhaopan, et al.
Published: (2025)
InstructPro: Natural Language Guided Ligand-Binding Protein Design
by: Song, Zhenqiao, et al.
Published: (2025)
by: Song, Zhenqiao, et al.
Published: (2025)
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
by: Meng, Fanqing, et al.
Published: (2024)
by: Meng, Fanqing, et al.
Published: (2024)
Edge-aware Decoding for Neural Asymmetric Routing
by: Liang, Li, et al.
Published: (2026)
by: Liang, Li, et al.
Published: (2026)
ProOPF: Benchmarking and Improving LLMs for Professional-Grade Power Systems Optimization Modeling
by: Shen, Chao, et al.
Published: (2026)
by: Shen, Chao, et al.
Published: (2026)
ForgeryGPT: A Multimodal LLM for Interpretable Image Forgery Detection and Localization
by: Zhang, Fanrui, et al.
Published: (2024)
by: Zhang, Fanrui, et al.
Published: (2024)
Story Arena: A Multi-Agent Environment for Envisioning the Future of Software Engineering
by: Weisz, Justin D., et al.
Published: (2025)
by: Weisz, Justin D., et al.
Published: (2025)
SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge
by: Li, Chuanhao, et al.
Published: (2024)
by: Li, Chuanhao, et al.
Published: (2024)
Yume-1.5: A Text-Controlled Interactive World Generation Model
by: Mao, Xiaofeng, et al.
Published: (2025)
by: Mao, Xiaofeng, et al.
Published: (2025)
DEVAL: A Framework for Evaluating and Improving the Derivation Capability of Large Language Models
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
Explore the Reasoning Capability of LLMs in the Chess Testbed
by: Wang, Shu, et al.
Published: (2024)
by: Wang, Shu, et al.
Published: (2024)
A Hierarchical and Attentional Analysis of Argument Structure Constructions in BERT Using Naturalistic Corpora
by: Kaipeng, Liu, et al.
Published: (2026)
by: Kaipeng, Liu, et al.
Published: (2026)
Multi-Step Reasoning for Embodied Question Answering via Tool Augmentation
by: Zhai, Mingliang, et al.
Published: (2025)
by: Zhai, Mingliang, et al.
Published: (2025)
Similar Items
-
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
by: Sun, Jianwen, et al.
Published: (2025) -
Closing the Expression Gap in LLM Instructions via Socratic Questioning
by: Sun, Jianwen, et al.
Published: (2025) -
From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration
by: Sun, Jianwen, et al.
Published: (2025) -
IA-T2I: Internet-Augmented Text-to-Image Generation
by: Li, Chuanhao, et al.
Published: (2025) -
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
by: Feng, Yukang, et al.
Published: (2025)