ARK: A Dual-Axis Multimodal Retrieval Benchmark along Reasoning and Knowledge

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lin, Yijie, Ding, Guofeng, Zhou, Haochen, Li, Haobin, Yang, Mouxing, Peng, Xi
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910017857257472
author Lin, Yijie
Ding, Guofeng
Zhou, Haochen
Li, Haobin
Yang, Mouxing
Peng, Xi
author_facet Lin, Yijie
Ding, Guofeng
Zhou, Haochen
Li, Haobin
Yang, Mouxing
Peng, Xi
contents Existing multimodal retrieval benchmarks largely emphasize semantic matching on daily-life images and offer limited diagnostics of professional knowledge and complex reasoning. To address this gap, we introduce ARK, a benchmark designed to analyze multimodal retrieval from two complementary perspectives: (i) knowledge domains (five domains with 17 subtypes), which characterize the content and expertise retrieval relies on, and (ii) reasoning skills (six categories), which characterize the type of inference over multimodal evidence required to identify the correct candidate. Specifically, ARK evaluates retrieval with both unimodal and multimodal queries and candidates, covering 16 heterogeneous visual data types. To avoid shortcut matching during evaluation, most queries are paired with targeted hard negatives that require multi-step reasoning. We evaluate 23 representative text-based and multimodal retrievers on ARK and observe a pronounced gap between knowledge-intensive and reasoning-intensive retrieval, with fine-grained visual and spatial reasoning emerging as persistent bottlenecks. We further show that simple enhancements such as re-ranking and rewriting yield consistent improvements, but substantial headroom remains.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09839
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ARK: A Dual-Axis Multimodal Retrieval Benchmark along Reasoning and Knowledge
Lin, Yijie
Ding, Guofeng
Zhou, Haochen
Li, Haobin
Yang, Mouxing
Peng, Xi
Computer Vision and Pattern Recognition
Existing multimodal retrieval benchmarks largely emphasize semantic matching on daily-life images and offer limited diagnostics of professional knowledge and complex reasoning. To address this gap, we introduce ARK, a benchmark designed to analyze multimodal retrieval from two complementary perspectives: (i) knowledge domains (five domains with 17 subtypes), which characterize the content and expertise retrieval relies on, and (ii) reasoning skills (six categories), which characterize the type of inference over multimodal evidence required to identify the correct candidate. Specifically, ARK evaluates retrieval with both unimodal and multimodal queries and candidates, covering 16 heterogeneous visual data types. To avoid shortcut matching during evaluation, most queries are paired with targeted hard negatives that require multi-step reasoning. We evaluate 23 representative text-based and multimodal retrievers on ARK and observe a pronounced gap between knowledge-intensive and reasoning-intensive retrieval, with fine-grained visual and spatial reasoning emerging as persistent bottlenecks. We further show that simple enhancements such as re-ranking and rewriting yield consistent improvements, but substantial headroom remains.
title ARK: A Dual-Axis Multimodal Retrieval Benchmark along Reasoning and Knowledge
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.09839