DocMMIR: A Framework for Document Multi-modal Information Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zirui, Wu, Siwei, Li, Yizhi, Wang, Xingyu, Zhou, Yi, Lin, Chenghua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914097406148608
author Li, Zirui
Wu, Siwei
Li, Yizhi
Wang, Xingyu
Zhou, Yi
Lin, Chenghua
author_facet Li, Zirui
Wu, Siwei
Li, Yizhi
Wang, Xingyu
Zhou, Yi
Lin, Chenghua
contents The rapid advancement of unsupervised representation learning and large-scale pre-trained vision-language models has significantly improved cross-modal retrieval tasks. However, existing multi-modal information retrieval (MMIR) studies lack a comprehensive exploration of document-level retrieval and suffer from the absence of cross-domain datasets at this granularity. To address this limitation, we introduce DocMMIR, a novel multi-modal document retrieval framework designed explicitly to unify diverse document formats and domains, including Wikipedia articles, scientific papers (arXiv), and presentation slides, within a comprehensive retrieval scenario. We construct a large-scale cross-domain multimodal benchmark, comprising 450K samples, which systematically integrates textual and visual information. Our comprehensive experimental analysis reveals substantial limitations in current state-of-the-art MLLMs (CLIP, BLIP2, SigLIP-2, ALIGN) when applied to our tasks, with only CLIP demonstrating reasonable zero-shot performance. Furthermore, we conduct a systematic investigation of training strategies, including cross-modal fusion methods and loss functions, and develop a tailored approach to train CLIP on our benchmark. This results in a +31% improvement in MRR@10 compared to the zero-shot baseline. All our data and code are released in https://github.com/J1mL1/DocMMIR.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19312
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DocMMIR: A Framework for Document Multi-modal Information Retrieval
Li, Zirui
Wu, Siwei
Li, Yizhi
Wang, Xingyu
Zhou, Yi
Lin, Chenghua
Information Retrieval
The rapid advancement of unsupervised representation learning and large-scale pre-trained vision-language models has significantly improved cross-modal retrieval tasks. However, existing multi-modal information retrieval (MMIR) studies lack a comprehensive exploration of document-level retrieval and suffer from the absence of cross-domain datasets at this granularity. To address this limitation, we introduce DocMMIR, a novel multi-modal document retrieval framework designed explicitly to unify diverse document formats and domains, including Wikipedia articles, scientific papers (arXiv), and presentation slides, within a comprehensive retrieval scenario. We construct a large-scale cross-domain multimodal benchmark, comprising 450K samples, which systematically integrates textual and visual information. Our comprehensive experimental analysis reveals substantial limitations in current state-of-the-art MLLMs (CLIP, BLIP2, SigLIP-2, ALIGN) when applied to our tasks, with only CLIP demonstrating reasonable zero-shot performance. Furthermore, we conduct a systematic investigation of training strategies, including cross-modal fusion methods and loss functions, and develop a tailored approach to train CLIP on our benchmark. This results in a +31% improvement in MRR@10 compared to the zero-shot baseline. All our data and code are released in https://github.com/J1mL1/DocMMIR.
title DocMMIR: A Framework for Document Multi-modal Information Retrieval
topic Information Retrieval
url https://arxiv.org/abs/2505.19312