MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866916905934127104 |
|---|---|
| author | Li, Shilong Bu, Xingyuan Wang, Wenjie Liu, Jiaheng Dong, Jun He, Haoyang Lu, Hao Zhang, Haozhe Jing, Chenchen Li, Zhen Li, Chuanhao Tian, Jiayi Zhang, Chenchen Peng, Tianhao He, Yancheng Gu, Jihao Zhang, Yuanxing Yang, Jian Zhang, Ge Huang, Wenhao Zhou, Wangchunshu Zhang, Zhaoxiang Ding, Ruizhe Wen, Shilei |
| author_facet | Li, Shilong Bu, Xingyuan Wang, Wenjie Liu, Jiaheng Dong, Jun He, Haoyang Lu, Hao Zhang, Haozhe Jing, Chenchen Li, Zhen Li, Chuanhao Tian, Jiayi Zhang, Chenchen Peng, Tianhao He, Yancheng Gu, Jihao Zhang, Yuanxing Yang, Jian Zhang, Ge Huang, Wenhao Zhou, Wangchunshu Zhang, Zhaoxiang Ding, Ruizhe Wen, Shilei |
| contents | AI agents with advanced reasoning and tool use capabilities have demonstrated impressive performance in web browsing for deep search. While existing benchmarks such as BrowseComp evaluate these browsing abilities, they primarily focus on textual information, overlooking the prevalence of multimodal content. To bridge this gap, we introduce MM-BrowseComp, a novel benchmark comprising 224 challenging, hand-crafted questions specifically designed to assess agents' multimodal retrieval and reasoning capabilities. These questions often incorporate images in prompts, and crucial information encountered during the search and reasoning process may also be embedded within images or videos on webpages. Consequently, methods relying solely on text prove insufficient for our benchmark. Additionally, we provide a verified checklist for each question, enabling fine-grained analysis of multimodal dependencies and reasoning paths. Our comprehensive evaluation of state-of-the-art models on MM-BrowseComp reveals that even top models like OpenAI o3 with tools achieve only 29.02\% accuracy, highlighting the suboptimal multimodal capabilities and lack of native multimodal reasoning in current models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_13186 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents Li, Shilong Bu, Xingyuan Wang, Wenjie Liu, Jiaheng Dong, Jun He, Haoyang Lu, Hao Zhang, Haozhe Jing, Chenchen Li, Zhen Li, Chuanhao Tian, Jiayi Zhang, Chenchen Peng, Tianhao He, Yancheng Gu, Jihao Zhang, Yuanxing Yang, Jian Zhang, Ge Huang, Wenhao Zhou, Wangchunshu Zhang, Zhaoxiang Ding, Ruizhe Wen, Shilei Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition AI agents with advanced reasoning and tool use capabilities have demonstrated impressive performance in web browsing for deep search. While existing benchmarks such as BrowseComp evaluate these browsing abilities, they primarily focus on textual information, overlooking the prevalence of multimodal content. To bridge this gap, we introduce MM-BrowseComp, a novel benchmark comprising 224 challenging, hand-crafted questions specifically designed to assess agents' multimodal retrieval and reasoning capabilities. These questions often incorporate images in prompts, and crucial information encountered during the search and reasoning process may also be embedded within images or videos on webpages. Consequently, methods relying solely on text prove insufficient for our benchmark. Additionally, we provide a verified checklist for each question, enabling fine-grained analysis of multimodal dependencies and reasoning paths. Our comprehensive evaluation of state-of-the-art models on MM-BrowseComp reveals that even top models like OpenAI o3 with tools achieve only 29.02\% accuracy, highlighting the suboptimal multimodal capabilities and lack of native multimodal reasoning in current models. |
| title | MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents |
| topic | Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2508.13186 |