InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hou, Bohan, Gu, Jiuning, Guo, Jiayan, Dang, Ronghao, Leng, Sicong, Li, Xin, Song, Xuemeng, Yang, Jianfei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915992286789632
author Hou, Bohan
Gu, Jiuning
Guo, Jiayan
Dang, Ronghao
Leng, Sicong
Li, Xin
Song, Xuemeng
Yang, Jianfei
author_facet Hou, Bohan
Gu, Jiuning
Guo, Jiayan
Dang, Ronghao
Leng, Sicong
Li, Xin
Song, Xuemeng
Yang, Jianfei
contents Existing benchmarks for multimodal agentic search evaluate multimodal search and visual browsing, but visual evidence is either confined to the input or treated as an answer endpoint rather than part of an interleaved search trajectory. We introduce \textbf{InterLV-Search}, a benchmark for Interleaved Language-Vision Agentic Search, in which textual and visual evidence is repeatedly used to condition later search. It contains 2,061 examples across three levels: active visual evidence seeking, controlled offline interleaved multimodal search, and open-web interleaved multimodal search. Beyond existing benchmarks, it also includes multimodal multi-branch samples that involve comparison between multiple entities during the evidence search. We construct Level 1 and Level 2 with automated pipelines and Level 3 with a machine-led, human-supervised open-web pipeline. We further provide InterLV-Agent for standardized tool use, trajectory logging, and evaluation. Experiments on proprietary and open-source multimodal agents show that current systems remain far from solving interleaved multimodal search, with the best model below 50% overall accuracy, highlighting challenges in visual evidence seeking, search control, and multimodal evidence integration. We release the benchmark data and evaluation code at https://github.com/hbhalpha/InterLV-Search-Bench
format Preprint
id arxiv_https___arxiv_org_abs_2605_07510
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search
Hou, Bohan
Gu, Jiuning
Guo, Jiayan
Dang, Ronghao
Leng, Sicong
Li, Xin
Song, Xuemeng
Yang, Jianfei
Computer Vision and Pattern Recognition
Computation and Language
Information Retrieval
Existing benchmarks for multimodal agentic search evaluate multimodal search and visual browsing, but visual evidence is either confined to the input or treated as an answer endpoint rather than part of an interleaved search trajectory. We introduce \textbf{InterLV-Search}, a benchmark for Interleaved Language-Vision Agentic Search, in which textual and visual evidence is repeatedly used to condition later search. It contains 2,061 examples across three levels: active visual evidence seeking, controlled offline interleaved multimodal search, and open-web interleaved multimodal search. Beyond existing benchmarks, it also includes multimodal multi-branch samples that involve comparison between multiple entities during the evidence search. We construct Level 1 and Level 2 with automated pipelines and Level 3 with a machine-led, human-supervised open-web pipeline. We further provide InterLV-Agent for standardized tool use, trajectory logging, and evaluation. Experiments on proprietary and open-source multimodal agents show that current systems remain far from solving interleaved multimodal search, with the best model below 50% overall accuracy, highlighting challenges in visual evidence seeking, search control, and multimodal evidence integration. We release the benchmark data and evaluation code at https://github.com/hbhalpha/InterLV-Search-Bench
title InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search
topic Computer Vision and Pattern Recognition
Computation and Language
Information Retrieval
url https://arxiv.org/abs/2605.07510