Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Dong, Kuicai, Chang, Yujing, Huang, Shijie, Wang, Yasheng, Tang, Ruiming, Liu, Yong
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911252759969792
author Dong, Kuicai
Chang, Yujing
Huang, Shijie
Wang, Yasheng
Tang, Ruiming
Liu, Yong
author_facet Dong, Kuicai
Chang, Yujing
Huang, Shijie
Wang, Yasheng
Tang, Ruiming
Liu, Yong
contents Document Visual Question Answering (DocVQA) faces dual challenges in processing lengthy multimodal documents (text, images, tables) and performing cross-modal reasoning. Current document retrieval-augmented generation (DocRAG) methods remain limited by their text-centric approaches, frequently missing critical visual information. The field also lacks robust benchmarks for assessing multimodal evidence selection and integration. We introduce MMDocRAG, a comprehensive benchmark featuring 4,055 expert-annotated QA pairs with multi-page, cross-modal evidence chains. Our framework introduces innovative metrics for evaluating multimodal quote selection and enables answers that interleave text with relevant visual elements. Through large-scale experiments with 60 VLM/LLM models and 14 retrieval systems, we identify persistent challenges in multimodal evidence retrieval, selection, and integration.Key findings reveal advanced proprietary LVMs show superior performance than open-sourced alternatives. Also, they show moderate advantages using multimodal inputs over text-only inputs, while open-source alternatives show significant performance degradation. Notably, fine-tuned LLMs achieve substantial improvements when using detailed image descriptions. MMDocRAG establishes a rigorous testing ground and provides actionable insights for developing more robust multimodal DocVQA systems. Our benchmark and code are available at https://mmdocrag.github.io/MMDocRAG/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16470
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
Dong, Kuicai
Chang, Yujing
Huang, Shijie
Wang, Yasheng
Tang, Ruiming
Liu, Yong
Information Retrieval
Computation and Language
Computer Vision and Pattern Recognition
Document Visual Question Answering (DocVQA) faces dual challenges in processing lengthy multimodal documents (text, images, tables) and performing cross-modal reasoning. Current document retrieval-augmented generation (DocRAG) methods remain limited by their text-centric approaches, frequently missing critical visual information. The field also lacks robust benchmarks for assessing multimodal evidence selection and integration. We introduce MMDocRAG, a comprehensive benchmark featuring 4,055 expert-annotated QA pairs with multi-page, cross-modal evidence chains. Our framework introduces innovative metrics for evaluating multimodal quote selection and enables answers that interleave text with relevant visual elements. Through large-scale experiments with 60 VLM/LLM models and 14 retrieval systems, we identify persistent challenges in multimodal evidence retrieval, selection, and integration.Key findings reveal advanced proprietary LVMs show superior performance than open-sourced alternatives. Also, they show moderate advantages using multimodal inputs over text-only inputs, while open-source alternatives show significant performance degradation. Notably, fine-tuned LLMs achieve substantial improvements when using detailed image descriptions. MMDocRAG establishes a rigorous testing ground and provides actionable insights for developing more robust multimodal DocVQA systems. Our benchmark and code are available at https://mmdocrag.github.io/MMDocRAG/.
title Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
topic Information Retrieval
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.16470