DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Narayan, Kartik, Xu, Yang, Cao, Tian, Nerella, Kavya, Patel, Vishal M., Shiee, Navid, Grasch, Peter, Jia, Chao, Yang, Yinfei, Gan, Zhe
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917013561016320
author Narayan, Kartik
Xu, Yang
Cao, Tian
Nerella, Kavya
Patel, Vishal M.
Shiee, Navid
Grasch, Peter
Jia, Chao
Yang, Yinfei
Gan, Zhe
author_facet Narayan, Kartik
Xu, Yang
Cao, Tian
Nerella, Kavya
Patel, Vishal M.
Shiee, Navid
Grasch, Peter
Jia, Chao
Yang, Yinfei
Gan, Zhe
contents Multimodal Large Language Models (MLLMs) in real-world applications require access to external knowledge sources and must remain responsive to the dynamic and ever-changing real-world information in order to address information-seeking and knowledge-intensive user queries. Existing approaches, such as retrieval augmented generation (RAG) methods, search agents, and search equipped MLLMs, often suffer from rigid pipelines, excessive search calls, and poorly constructed search queries, which result in inefficiencies and suboptimal outcomes. To address these limitations, we present DeepMMSearch-R1, the first multimodal LLM capable of performing on-demand, multi-turn web searches and dynamically crafting queries for both image and text search tools. Specifically, DeepMMSearch-R1 can initiate web searches based on relevant crops of the input image making the image search more effective, and can iteratively adapt text search queries based on retrieved information, thereby enabling self-reflection and self-correction. Our approach relies on a two-stage training pipeline: a cold start supervised finetuning phase followed by an online reinforcement learning optimization. For training, we introduce DeepMMSearchVQA, a novel multimodal VQA dataset created through an automated pipeline intermixed with real-world information from web search tools. This dataset contains diverse, multi-hop queries that integrate textual and visual information, teaching the model when to search, what to search for, which search tool to use and how to reason over the retrieved information. We conduct extensive experiments across a range of knowledge-intensive benchmarks to demonstrate the superiority of our approach. Finally, we analyze the results and provide insights that are valuable for advancing multimodal web-search.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12801
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search
Narayan, Kartik
Xu, Yang
Cao, Tian
Nerella, Kavya
Patel, Vishal M.
Shiee, Navid
Grasch, Peter
Jia, Chao
Yang, Yinfei
Gan, Zhe
Computer Vision and Pattern Recognition
Information Retrieval
Multimodal Large Language Models (MLLMs) in real-world applications require access to external knowledge sources and must remain responsive to the dynamic and ever-changing real-world information in order to address information-seeking and knowledge-intensive user queries. Existing approaches, such as retrieval augmented generation (RAG) methods, search agents, and search equipped MLLMs, often suffer from rigid pipelines, excessive search calls, and poorly constructed search queries, which result in inefficiencies and suboptimal outcomes. To address these limitations, we present DeepMMSearch-R1, the first multimodal LLM capable of performing on-demand, multi-turn web searches and dynamically crafting queries for both image and text search tools. Specifically, DeepMMSearch-R1 can initiate web searches based on relevant crops of the input image making the image search more effective, and can iteratively adapt text search queries based on retrieved information, thereby enabling self-reflection and self-correction. Our approach relies on a two-stage training pipeline: a cold start supervised finetuning phase followed by an online reinforcement learning optimization. For training, we introduce DeepMMSearchVQA, a novel multimodal VQA dataset created through an automated pipeline intermixed with real-world information from web search tools. This dataset contains diverse, multi-hop queries that integrate textual and visual information, teaching the model when to search, what to search for, which search tool to use and how to reason over the retrieved information. We conduct extensive experiments across a range of knowledge-intensive benchmarks to demonstrate the superiority of our approach. Finally, we analyze the results and provide insights that are valuable for advancing multimodal web-search.
title DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search
topic Computer Vision and Pattern Recognition
Information Retrieval
url https://arxiv.org/abs/2510.12801