Inference-Free Multimodal Learned Sparse Retrieval for Production-Scale Visual Document Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cho, Gyu-Hwung, Lee, Youngjune, Jeong, Kiyoon, Lee, Siyoung, Han, Sanggyu, Dejean, Hervé, Clinchant, Stéphane, Hwang, Seung-won
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910271800344576
author Cho, Gyu-Hwung
Lee, Youngjune
Jeong, Kiyoon
Lee, Siyoung
Han, Sanggyu
Dejean, Hervé
Clinchant, Stéphane
Hwang, Seung-won
author_facet Cho, Gyu-Hwung
Lee, Youngjune
Jeong, Kiyoon
Lee, Siyoung
Han, Sanggyu
Dejean, Hervé
Clinchant, Stéphane
Hwang, Seung-won
contents As large-scale visual-document corpora such as arXiv papers and enterprise PDFs continue to grow, visual-document retrieval has gained increasing attention; yet it still lacks a deployable system that lexically indexes visual documents to serve queries without neural encoding at scale. Existing methods either achieve strong retrieval quality with VLM-based dense or multi-vector models but require neural query encoding at serving time, or avoid query encoding with OCR- or caption-based BM25 at the cost of time-consuming text extraction or generation. To fill this missing serving regime, we present V-SPLADE, an inference-free sparse retriever for visual-document retrieval. However, such inference-free multimodal learned sparse retrieval systems remain underexplored and have not yet shown dense-level effectiveness under high sparsity. We attribute this limitation to a lexical grounding problem: visual sparse representations often fail to capture the lexical content embedded in document images. To address this problem, we introduce caption-gated token supervision, a training-only signal that uses VLM-generated captions as lexical cues to activate retrieval-relevant vocabulary dimensions. With this supervision, V-SPLADE improves average NDCG@5 across six visual-document retrieval benchmarks by +13.8pp over the same-scale dense baseline and by up to +6.3pp over OCR- or caption-based BM25 baselines. On an 18.7M-document corpus, it more than doubles R@5 over the same-scale dense baseline and further improves competing retrievers through score fusion by up to +2.4pp R@5. Code will be released soon at https://github.com/naver/v-splade.
format Preprint
id arxiv_https___arxiv_org_abs_2605_30917
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Inference-Free Multimodal Learned Sparse Retrieval for Production-Scale Visual Document Search
Cho, Gyu-Hwung
Lee, Youngjune
Jeong, Kiyoon
Lee, Siyoung
Han, Sanggyu
Dejean, Hervé
Clinchant, Stéphane
Hwang, Seung-won
Information Retrieval
Computer Vision and Pattern Recognition
H.3.3; I.2.7
As large-scale visual-document corpora such as arXiv papers and enterprise PDFs continue to grow, visual-document retrieval has gained increasing attention; yet it still lacks a deployable system that lexically indexes visual documents to serve queries without neural encoding at scale. Existing methods either achieve strong retrieval quality with VLM-based dense or multi-vector models but require neural query encoding at serving time, or avoid query encoding with OCR- or caption-based BM25 at the cost of time-consuming text extraction or generation. To fill this missing serving regime, we present V-SPLADE, an inference-free sparse retriever for visual-document retrieval. However, such inference-free multimodal learned sparse retrieval systems remain underexplored and have not yet shown dense-level effectiveness under high sparsity. We attribute this limitation to a lexical grounding problem: visual sparse representations often fail to capture the lexical content embedded in document images. To address this problem, we introduce caption-gated token supervision, a training-only signal that uses VLM-generated captions as lexical cues to activate retrieval-relevant vocabulary dimensions. With this supervision, V-SPLADE improves average NDCG@5 across six visual-document retrieval benchmarks by +13.8pp over the same-scale dense baseline and by up to +6.3pp over OCR- or caption-based BM25 baselines. On an 18.7M-document corpus, it more than doubles R@5 over the same-scale dense baseline and further improves competing retrievers through score fusion by up to +2.4pp R@5. Code will be released soon at https://github.com/naver/v-splade.
title Inference-Free Multimodal Learned Sparse Retrieval for Production-Scale Visual Document Search
topic Information Retrieval
Computer Vision and Pattern Recognition
H.3.3; I.2.7
url https://arxiv.org/abs/2605.30917