Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zeng, Yu, Huang, Wenxuan, Fang, Zhen, Chen, Shuang, Shen, Yufan, Cai, Yishuo, Wang, Xiaoman, Yin, Zhenfei, Chen, Lin, Chen, Zehui, Huang, Shiting, Zhao, Yiming, Tang, Xu, Hu, Yao, Torr, Philip, Ouyang, Wanli, Cao, Shaosheng
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908858250690560
author Zeng, Yu
Huang, Wenxuan
Fang, Zhen
Chen, Shuang
Shen, Yufan
Cai, Yishuo
Wang, Xiaoman
Yin, Zhenfei
Chen, Lin
Chen, Zehui
Huang, Shiting
Zhao, Yiming
Tang, Xu
Hu, Yao
Torr, Philip
Ouyang, Wanli
Cao, Shaosheng
author_facet Zeng, Yu
Huang, Wenxuan
Fang, Zhen
Chen, Shuang
Shen, Yufan
Cai, Yishuo
Wang, Xiaoman
Yin, Zhenfei
Chen, Lin
Chen, Zehui
Huang, Shiting
Zhao, Yiming
Tang, Xu
Hu, Yao
Torr, Philip
Ouyang, Wanli
Cao, Shaosheng
contents Multimodal Large Language Models (MLLMs) have advanced VQA and now support Vision-DeepResearch systems that use search engines for complex visual-textual fact-finding. However, evaluating these visual and textual search abilities is still difficult, and existing benchmarks have two major limitations. First, existing benchmarks are not visual search-centric: answers that should require visual search are often leaked through cross-textual cues in the text questions or can be inferred from the prior world knowledge in current MLLMs. Second, overly idealized evaluation scenario: On the image-search side, the required information can often be obtained via near-exact matching against the full image, while the text-search side is overly direct and insufficiently challenging. To address these issues, we construct the Vision-DeepResearch benchmark (VDR-Bench) comprising 2,000 VQA instances. All questions are created via a careful, multi-stage curation pipeline and rigorous expert review, designed to assess the behavior of Vision-DeepResearch systems under realistic real-world conditions. Moreover, to address the insufficient visual retrieval capabilities of current MLLMs, we propose a simple multi-round cropped-search workflow. This strategy is shown to effectively improve model performance in realistic visual retrieval scenarios. Overall, our results provide practical guidance for the design of future multimodal deep-research systems. The code will be released in https://github.com/Osilly/Vision-DeepResearch.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02185
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models
Zeng, Yu
Huang, Wenxuan
Fang, Zhen
Chen, Shuang
Shen, Yufan
Cai, Yishuo
Wang, Xiaoman
Yin, Zhenfei
Chen, Lin
Chen, Zehui
Huang, Shiting
Zhao, Yiming
Tang, Xu
Hu, Yao
Torr, Philip
Ouyang, Wanli
Cao, Shaosheng
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Multimodal Large Language Models (MLLMs) have advanced VQA and now support Vision-DeepResearch systems that use search engines for complex visual-textual fact-finding. However, evaluating these visual and textual search abilities is still difficult, and existing benchmarks have two major limitations. First, existing benchmarks are not visual search-centric: answers that should require visual search are often leaked through cross-textual cues in the text questions or can be inferred from the prior world knowledge in current MLLMs. Second, overly idealized evaluation scenario: On the image-search side, the required information can often be obtained via near-exact matching against the full image, while the text-search side is overly direct and insufficiently challenging. To address these issues, we construct the Vision-DeepResearch benchmark (VDR-Bench) comprising 2,000 VQA instances. All questions are created via a careful, multi-stage curation pipeline and rigorous expert review, designed to assess the behavior of Vision-DeepResearch systems under realistic real-world conditions. Moreover, to address the insufficient visual retrieval capabilities of current MLLMs, we propose a simple multi-round cropped-search workflow. This strategy is shown to effectively improve model performance in realistic visual retrieval scenarios. Overall, our results provide practical guidance for the design of future multimodal deep-research systems. The code will be released in https://github.com/Osilly/Vision-DeepResearch.
title Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2602.02185