MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Shilong, Bu, Xingyuan, Wang, Wenjie, Liu, Jiaheng, Dong, Jun, He, Haoyang, Lu, Hao, Zhang, Haozhe, Jing, Chenchen, Li, Zhen, Li, Chuanhao, Tian, Jiayi, Zhang, Chenchen, Peng, Tianhao, He, Yancheng, Gu, Jihao, Zhang, Yuanxing, Yang, Jian, Zhang, Ge, Huang, Wenhao, Zhou, Wangchunshu, Zhang, Zhaoxiang, Ding, Ruizhe, Wen, Shilei
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916905934127104
author Li, Shilong
Bu, Xingyuan
Wang, Wenjie
Liu, Jiaheng
Dong, Jun
He, Haoyang
Lu, Hao
Zhang, Haozhe
Jing, Chenchen
Li, Zhen
Li, Chuanhao
Tian, Jiayi
Zhang, Chenchen
Peng, Tianhao
He, Yancheng
Gu, Jihao
Zhang, Yuanxing
Yang, Jian
Zhang, Ge
Huang, Wenhao
Zhou, Wangchunshu
Zhang, Zhaoxiang
Ding, Ruizhe
Wen, Shilei
author_facet Li, Shilong
Bu, Xingyuan
Wang, Wenjie
Liu, Jiaheng
Dong, Jun
He, Haoyang
Lu, Hao
Zhang, Haozhe
Jing, Chenchen
Li, Zhen
Li, Chuanhao
Tian, Jiayi
Zhang, Chenchen
Peng, Tianhao
He, Yancheng
Gu, Jihao
Zhang, Yuanxing
Yang, Jian
Zhang, Ge
Huang, Wenhao
Zhou, Wangchunshu
Zhang, Zhaoxiang
Ding, Ruizhe
Wen, Shilei
contents AI agents with advanced reasoning and tool use capabilities have demonstrated impressive performance in web browsing for deep search. While existing benchmarks such as BrowseComp evaluate these browsing abilities, they primarily focus on textual information, overlooking the prevalence of multimodal content. To bridge this gap, we introduce MM-BrowseComp, a novel benchmark comprising 224 challenging, hand-crafted questions specifically designed to assess agents' multimodal retrieval and reasoning capabilities. These questions often incorporate images in prompts, and crucial information encountered during the search and reasoning process may also be embedded within images or videos on webpages. Consequently, methods relying solely on text prove insufficient for our benchmark. Additionally, we provide a verified checklist for each question, enabling fine-grained analysis of multimodal dependencies and reasoning paths. Our comprehensive evaluation of state-of-the-art models on MM-BrowseComp reveals that even top models like OpenAI o3 with tools achieve only 29.02\% accuracy, highlighting the suboptimal multimodal capabilities and lack of native multimodal reasoning in current models.
format Preprint
id arxiv_https___arxiv_org_abs_2508_13186
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
Li, Shilong
Bu, Xingyuan
Wang, Wenjie
Liu, Jiaheng
Dong, Jun
He, Haoyang
Lu, Hao
Zhang, Haozhe
Jing, Chenchen
Li, Zhen
Li, Chuanhao
Tian, Jiayi
Zhang, Chenchen
Peng, Tianhao
He, Yancheng
Gu, Jihao
Zhang, Yuanxing
Yang, Jian
Zhang, Ge
Huang, Wenhao
Zhou, Wangchunshu
Zhang, Zhaoxiang
Ding, Ruizhe
Wen, Shilei
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
AI agents with advanced reasoning and tool use capabilities have demonstrated impressive performance in web browsing for deep search. While existing benchmarks such as BrowseComp evaluate these browsing abilities, they primarily focus on textual information, overlooking the prevalence of multimodal content. To bridge this gap, we introduce MM-BrowseComp, a novel benchmark comprising 224 challenging, hand-crafted questions specifically designed to assess agents' multimodal retrieval and reasoning capabilities. These questions often incorporate images in prompts, and crucial information encountered during the search and reasoning process may also be embedded within images or videos on webpages. Consequently, methods relying solely on text prove insufficient for our benchmark. Additionally, we provide a verified checklist for each question, enabling fine-grained analysis of multimodal dependencies and reasoning paths. Our comprehensive evaluation of state-of-the-art models on MM-BrowseComp reveals that even top models like OpenAI o3 with tools achieve only 29.02\% accuracy, highlighting the suboptimal multimodal capabilities and lack of native multimodal reasoning in current models.
title MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.13186