Unifying Multimodal Retrieval via Document Screenshot Embedding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Xueguang, Lin, Sheng-Chieh, Li, Minghan, Chen, Wenhu, Lin, Jimmy
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912140136284160
author Ma, Xueguang
Lin, Sheng-Chieh
Li, Minghan
Chen, Wenhu
Lin, Jimmy
author_facet Ma, Xueguang
Lin, Sheng-Chieh
Li, Minghan
Chen, Wenhu
Lin, Jimmy
contents In the real world, documents are organized in different formats and varied modalities. Traditional retrieval pipelines require tailored document parsing techniques and content extraction modules to prepare input for indexing. This process is tedious, prone to errors, and has information loss. To this end, we propose Document Screenshot Embedding (DSE), a novel retrieval paradigm that regards document screenshots as a unified input format, which does not require any content extraction preprocess and preserves all the information in a document (e.g., text, image and layout). DSE leverages a large vision-language model to directly encode document screenshots into dense representations for retrieval. To evaluate our method, we first craft the dataset of Wiki-SS, a 1.3M Wikipedia web page screenshots as the corpus to answer the questions from the Natural Questions dataset. In such a text-intensive document retrieval setting, DSE shows competitive effectiveness compared to other text retrieval methods relying on parsing. For example, DSE outperforms BM25 by 17 points in top-1 retrieval accuracy. Additionally, in a mixed-modality task of slide retrieval, DSE significantly outperforms OCR text retrieval methods by over 15 points in nDCG@10. These experiments show that DSE is an effective document retrieval paradigm for diverse types of documents. Model checkpoints, code, and Wiki-SS collection will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2406_11251
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unifying Multimodal Retrieval via Document Screenshot Embedding
Ma, Xueguang
Lin, Sheng-Chieh
Li, Minghan
Chen, Wenhu
Lin, Jimmy
Information Retrieval
In the real world, documents are organized in different formats and varied modalities. Traditional retrieval pipelines require tailored document parsing techniques and content extraction modules to prepare input for indexing. This process is tedious, prone to errors, and has information loss. To this end, we propose Document Screenshot Embedding (DSE), a novel retrieval paradigm that regards document screenshots as a unified input format, which does not require any content extraction preprocess and preserves all the information in a document (e.g., text, image and layout). DSE leverages a large vision-language model to directly encode document screenshots into dense representations for retrieval. To evaluate our method, we first craft the dataset of Wiki-SS, a 1.3M Wikipedia web page screenshots as the corpus to answer the questions from the Natural Questions dataset. In such a text-intensive document retrieval setting, DSE shows competitive effectiveness compared to other text retrieval methods relying on parsing. For example, DSE outperforms BM25 by 17 points in top-1 retrieval accuracy. Additionally, in a mixed-modality task of slide retrieval, DSE significantly outperforms OCR text retrieval methods by over 15 points in nDCG@10. These experiments show that DSE is an effective document retrieval paradigm for diverse types of documents. Model checkpoints, code, and Wiki-SS collection will be released.
title Unifying Multimodal Retrieval via Document Screenshot Embedding
topic Information Retrieval
url https://arxiv.org/abs/2406.11251