MARVEL: Unlocking the Multi-Modal Capability of Dense Retrieval via Visual Module Plugin

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Tianshuo, Mei, Sen, Li, Xinze, Liu, Zhenghao, Xiong, Chenyan, Liu, Zhiyuan, Gu, Yu, Yu, Ge
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914834823512064
author Zhou, Tianshuo
Mei, Sen
Li, Xinze
Liu, Zhenghao
Xiong, Chenyan
Liu, Zhiyuan
Gu, Yu
Yu, Ge
author_facet Zhou, Tianshuo
Mei, Sen
Li, Xinze
Liu, Zhenghao
Xiong, Chenyan
Liu, Zhiyuan
Gu, Yu
Yu, Ge
contents This paper proposes Multi-modAl Retrieval model via Visual modulE pLugin (MARVEL), which learns an embedding space for queries and multi-modal documents to conduct retrieval. MARVEL encodes queries and multi-modal documents with a unified encoder model, which helps to alleviate the modality gap between images and texts. Specifically, we enable the image understanding ability of the well-trained dense retriever, T5-ANCE, by incorporating the visual module's encoded image features as its inputs. To facilitate the multi-modal retrieval tasks, we build the ClueWeb22-MM dataset based on the ClueWeb22 dataset, which regards anchor texts as queries, and extracts the related text and image documents from anchor-linked web pages. Our experiments show that MARVEL significantly outperforms the state-of-the-art methods on the multi-modal retrieval dataset WebQA and ClueWeb22-MM. MARVEL provides an opportunity to broaden the advantages of text retrieval to the multi-modal scenario. Besides, we also illustrate that the language model has the ability to extract image semantics and partly map the image features to the input word embedding space. All codes are available at https://github.com/OpenMatch/MARVEL.
format Preprint
id arxiv_https___arxiv_org_abs_2310_14037
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MARVEL: Unlocking the Multi-Modal Capability of Dense Retrieval via Visual Module Plugin
Zhou, Tianshuo
Mei, Sen
Li, Xinze
Liu, Zhenghao
Xiong, Chenyan
Liu, Zhiyuan
Gu, Yu
Yu, Ge
Information Retrieval
This paper proposes Multi-modAl Retrieval model via Visual modulE pLugin (MARVEL), which learns an embedding space for queries and multi-modal documents to conduct retrieval. MARVEL encodes queries and multi-modal documents with a unified encoder model, which helps to alleviate the modality gap between images and texts. Specifically, we enable the image understanding ability of the well-trained dense retriever, T5-ANCE, by incorporating the visual module's encoded image features as its inputs. To facilitate the multi-modal retrieval tasks, we build the ClueWeb22-MM dataset based on the ClueWeb22 dataset, which regards anchor texts as queries, and extracts the related text and image documents from anchor-linked web pages. Our experiments show that MARVEL significantly outperforms the state-of-the-art methods on the multi-modal retrieval dataset WebQA and ClueWeb22-MM. MARVEL provides an opportunity to broaden the advantages of text retrieval to the multi-modal scenario. Besides, we also illustrate that the language model has the ability to extract image semantics and partly map the image features to the input word embedding space. All codes are available at https://github.com/OpenMatch/MARVEL.
title MARVEL: Unlocking the Multi-Modal Capability of Dense Retrieval via Visual Module Plugin
topic Information Retrieval
url https://arxiv.org/abs/2310.14037