MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Hui, Zhai, Haoquan, Li, Yuchen, Cai, Hengyi, Zhang, Peirong, Zhang, Yidan, Wang, Lei, Wang, Chunle, Hou, Yingyan, Wang, Shuaiqiang, Yin, Dawei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913041579245568
author Wu, Hui
Zhai, Haoquan
Li, Yuchen
Cai, Hengyi
Zhang, Peirong
Zhang, Yidan
Wang, Lei
Wang, Chunle
Hou, Yingyan
Wang, Shuaiqiang
Yin, Dawei
author_facet Wu, Hui
Zhai, Haoquan
Li, Yuchen
Cai, Hengyi
Zhang, Peirong
Zhang, Yidan
Wang, Lei
Wang, Chunle
Hou, Yingyan
Wang, Shuaiqiang
Yin, Dawei
contents Retrieval-based multimodal document QA aims to identify and integrate relevant information from visually rich documents with complex multimodal structures. While retrieval-augmented generation (RAG) has shown strong performance in text-based QA, its extensions to multimodal documents remain underexplored and face significant limitations. Specifically, current approaches rely on query-agnostic document representations that overlook salient content and use static top-k evidence selection, which fails to adapt to the uncertain distribution of relevant information. To address these limitations, we propose the Multimodal Adaptive Retrieval-Augmented (MARA) framework, which introduces query-adaptive mechanisms to both retrieval and generation. MARA consists of two components: a Query-Aligned Region Encoder that builds multi-level document representations and reweights them based on query relevance to improve retrieval precision; and a Self-Reflective Evidence Controller that monitors evidence sufficiency during generation and adaptively incorporates content from lower-ranked sources using a sliding-window strategy. Experiments on six multimodal QA benchmarks demonstrate that MARA consistently improves retrieval relevance and answer quality over existing SOTA method.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16313
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering
Wu, Hui
Zhai, Haoquan
Li, Yuchen
Cai, Hengyi
Zhang, Peirong
Zhang, Yidan
Wang, Lei
Wang, Chunle
Hou, Yingyan
Wang, Shuaiqiang
Yin, Dawei
Information Retrieval
Artificial Intelligence
Computation and Language
Retrieval-based multimodal document QA aims to identify and integrate relevant information from visually rich documents with complex multimodal structures. While retrieval-augmented generation (RAG) has shown strong performance in text-based QA, its extensions to multimodal documents remain underexplored and face significant limitations. Specifically, current approaches rely on query-agnostic document representations that overlook salient content and use static top-k evidence selection, which fails to adapt to the uncertain distribution of relevant information. To address these limitations, we propose the Multimodal Adaptive Retrieval-Augmented (MARA) framework, which introduces query-adaptive mechanisms to both retrieval and generation. MARA consists of two components: a Query-Aligned Region Encoder that builds multi-level document representations and reweights them based on query relevance to improve retrieval precision; and a Self-Reflective Evidence Controller that monitors evidence sufficiency during generation and adaptively incorporates content from lower-ranked sources using a sliding-window strategy. Experiments on six multimodal QA benchmarks demonstrate that MARA consistently improves retrieval relevance and answer quality over existing SOTA method.
title MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2604.16313