Enhancing Multimodal Retrieval via Complementary Information Extraction and Alignment

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zeng, Delong, Xie, Yuexiang, Li, Yaliang, Shen, Ying
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915715687120896
author Zeng, Delong
Xie, Yuexiang
Li, Yaliang
Shen, Ying
author_facet Zeng, Delong
Xie, Yuexiang
Li, Yaliang
Shen, Ying
contents Multimodal retrieval has emerged as a promising yet challenging research direction in recent years. Most existing studies in multimodal retrieval focus on capturing information in multimodal data that is similar to their paired texts, but often ignores the complementary information contained in multimodal data. In this study, we propose CIEA, a novel multimodal retrieval approach that employs Complementary Information Extraction and Alignment, which transforms both text and images in documents into a unified latent space and features a complementary information extractor designed to identify and preserve differences in the image representations. We optimize CIEA using two complementary contrastive losses to ensure semantic integrity and effectively capture the complementary information contained in images. Extensive experiments demonstrate the effectiveness of CIEA, which achieves significant improvements over both divide-and-conquer models and universal dense retrieval models. We provide an ablation study, further discussions, and case studies to highlight the advancements achieved by CIEA. To promote further research in the community, we have released the source code at https://github.com/zengdlong/CIEA.
format Preprint
id arxiv_https___arxiv_org_abs_2601_04571
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Enhancing Multimodal Retrieval via Complementary Information Extraction and Alignment
Zeng, Delong
Xie, Yuexiang
Li, Yaliang
Shen, Ying
Artificial Intelligence
Multimedia
Multimodal retrieval has emerged as a promising yet challenging research direction in recent years. Most existing studies in multimodal retrieval focus on capturing information in multimodal data that is similar to their paired texts, but often ignores the complementary information contained in multimodal data. In this study, we propose CIEA, a novel multimodal retrieval approach that employs Complementary Information Extraction and Alignment, which transforms both text and images in documents into a unified latent space and features a complementary information extractor designed to identify and preserve differences in the image representations. We optimize CIEA using two complementary contrastive losses to ensure semantic integrity and effectively capture the complementary information contained in images. Extensive experiments demonstrate the effectiveness of CIEA, which achieves significant improvements over both divide-and-conquer models and universal dense retrieval models. We provide an ablation study, further discussions, and case studies to highlight the advancements achieved by CIEA. To promote further research in the community, we have released the source code at https://github.com/zengdlong/CIEA.
title Enhancing Multimodal Retrieval via Complementary Information Extraction and Alignment
topic Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2601.04571