ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yan, Li, Dongxu, Wu, Haoning, Chen, Bei, Liu, Liu, Pan, Liyuan, Li, Junnan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915189614444544
author Yang, Yan
Li, Dongxu
Wu, Haoning
Chen, Bei
Liu, Liu
Pan, Liyuan
Li, Junnan
author_facet Yang, Yan
Li, Dongxu
Wu, Haoning
Chen, Bei
Liu, Liu
Pan, Liyuan
Li, Junnan
contents Solving expert-level multimodal tasks is a key milestone towards general intelligence. As the capabilities of multimodal large language models (MLLMs) continue to improve, evaluation of such advanced multimodal intelligence becomes necessary yet challenging. In this work, we introduce ProBench, a benchmark of open-ended user queries that require professional expertise and advanced reasoning. ProBench consists of 4,000 high-quality samples independently submitted by professionals based on their daily productivity demands. It spans across 10 fields and 56 sub-fields, including science, arts, humanities, coding, mathematics, and creative writing. Experimentally, we evaluate and compare 24 latest models using MLLM-as-a-Judge. Our results reveal that although the best open-source models rival the proprietary ones, ProBench presents significant challenges in visual perception, textual understanding, domain knowledge and advanced reasoning, thus providing valuable directions for future multimodal AI research efforts.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06885
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks
Yang, Yan
Li, Dongxu
Wu, Haoning
Chen, Bei
Liu, Liu
Pan, Liyuan
Li, Junnan
Computer Vision and Pattern Recognition
Solving expert-level multimodal tasks is a key milestone towards general intelligence. As the capabilities of multimodal large language models (MLLMs) continue to improve, evaluation of such advanced multimodal intelligence becomes necessary yet challenging. In this work, we introduce ProBench, a benchmark of open-ended user queries that require professional expertise and advanced reasoning. ProBench consists of 4,000 high-quality samples independently submitted by professionals based on their daily productivity demands. It spans across 10 fields and 56 sub-fields, including science, arts, humanities, coding, mathematics, and creative writing. Experimentally, we evaluate and compare 24 latest models using MLLM-as-a-Judge. Our results reveal that although the best open-source models rival the proprietary ones, ProBench presents significant challenges in visual perception, textual understanding, domain knowledge and advanced reasoning, thus providing valuable directions for future multimodal AI research efforts.
title ProBench: Judging Multimodal Foundation Models on Open-ended Multi-domain Expert Tasks
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.06885