Document-as-Image Representations Fall Short for Scientific Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khalighinejad, Ghazal, Thirukovalluru, Raghuveer, Oh, Alexander H., Dhingra, Bhuwan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915945978527744
author Khalighinejad, Ghazal
Thirukovalluru, Raghuveer
Oh, Alexander H.
Dhingra, Bhuwan
author_facet Khalighinejad, Ghazal
Thirukovalluru, Raghuveer
Oh, Alexander H.
Dhingra, Bhuwan
contents Many recent document embedding models are trained on document-as-image representations, embedding rendered pages as images rather than the underlying source. Meanwhile, existing benchmarks for scientific document retrieval, such as ArXivQA and ViDoRe, treat documents as images of pages, implicitly favoring such representations. In this work, we argue that this paradigm is not well-suited for text-rich multimodal scientific documents, where critical evidence is distributed across structured sources, including text, tables, and figures. To study this setting, we introduce ArXivDoc, a new benchmark constructed from the underlying LaTeX sources of scientific papers. Unlike PDF or image-based representations, LaTeX provides direct access to structured elements (e.g., sections, tables, figures, equations), enabling controlled query construction grounded in specific evidence types. We systematically compare text-only, image-based, and multimodal representations across both single-vector and multi-vector retrieval models. Our results show that: (1) document-as-image representations are consistently suboptimal, especially as document length increases; (2) text-based representations are most effective, even for figure-based queries, by leveraging captions and surrounding context; and (3) interleaved text+image representations outperform document-as-image approaches without requiring specialized training.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18508
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Document-as-Image Representations Fall Short for Scientific Retrieval
Khalighinejad, Ghazal
Thirukovalluru, Raghuveer
Oh, Alexander H.
Dhingra, Bhuwan
Information Retrieval
Artificial Intelligence
Computation and Language
Many recent document embedding models are trained on document-as-image representations, embedding rendered pages as images rather than the underlying source. Meanwhile, existing benchmarks for scientific document retrieval, such as ArXivQA and ViDoRe, treat documents as images of pages, implicitly favoring such representations. In this work, we argue that this paradigm is not well-suited for text-rich multimodal scientific documents, where critical evidence is distributed across structured sources, including text, tables, and figures. To study this setting, we introduce ArXivDoc, a new benchmark constructed from the underlying LaTeX sources of scientific papers. Unlike PDF or image-based representations, LaTeX provides direct access to structured elements (e.g., sections, tables, figures, equations), enabling controlled query construction grounded in specific evidence types. We systematically compare text-only, image-based, and multimodal representations across both single-vector and multi-vector retrieval models. Our results show that: (1) document-as-image representations are consistently suboptimal, especially as document length increases; (2) text-based representations are most effective, even for figure-based queries, by leveraging captions and surrounding context; and (3) interleaved text+image representations outperform document-as-image approaches without requiring specialized training.
title Document-as-Image Representations Fall Short for Scientific Retrieval
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2604.18508