MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913041579245568 |
|---|---|
| author | Wu, Hui Zhai, Haoquan Li, Yuchen Cai, Hengyi Zhang, Peirong Zhang, Yidan Wang, Lei Wang, Chunle Hou, Yingyan Wang, Shuaiqiang Yin, Dawei |
| author_facet | Wu, Hui Zhai, Haoquan Li, Yuchen Cai, Hengyi Zhang, Peirong Zhang, Yidan Wang, Lei Wang, Chunle Hou, Yingyan Wang, Shuaiqiang Yin, Dawei |
| contents | Retrieval-based multimodal document QA aims to identify and integrate relevant information from visually rich documents with complex multimodal structures. While retrieval-augmented generation (RAG) has shown strong performance in text-based QA, its extensions to multimodal documents remain underexplored and face significant limitations. Specifically, current approaches rely on query-agnostic document representations that overlook salient content and use static top-k evidence selection, which fails to adapt to the uncertain distribution of relevant information. To address these limitations, we propose the Multimodal Adaptive Retrieval-Augmented (MARA) framework, which introduces query-adaptive mechanisms to both retrieval and generation. MARA consists of two components: a Query-Aligned Region Encoder that builds multi-level document representations and reweights them based on query relevance to improve retrieval precision; and a Self-Reflective Evidence Controller that monitors evidence sufficiency during generation and adaptively incorporates content from lower-ranked sources using a sliding-window strategy. Experiments on six multimodal QA benchmarks demonstrate that MARA consistently improves retrieval relevance and answer quality over existing SOTA method. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_16313 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering Wu, Hui Zhai, Haoquan Li, Yuchen Cai, Hengyi Zhang, Peirong Zhang, Yidan Wang, Lei Wang, Chunle Hou, Yingyan Wang, Shuaiqiang Yin, Dawei Information Retrieval Artificial Intelligence Computation and Language Retrieval-based multimodal document QA aims to identify and integrate relevant information from visually rich documents with complex multimodal structures. While retrieval-augmented generation (RAG) has shown strong performance in text-based QA, its extensions to multimodal documents remain underexplored and face significant limitations. Specifically, current approaches rely on query-agnostic document representations that overlook salient content and use static top-k evidence selection, which fails to adapt to the uncertain distribution of relevant information. To address these limitations, we propose the Multimodal Adaptive Retrieval-Augmented (MARA) framework, which introduces query-adaptive mechanisms to both retrieval and generation. MARA consists of two components: a Query-Aligned Region Encoder that builds multi-level document representations and reweights them based on query relevance to improve retrieval precision; and a Self-Reflective Evidence Controller that monitors evidence sufficiency during generation and adaptively incorporates content from lower-ranked sources using a sliding-window strategy. Experiments on six multimodal QA benchmarks demonstrate that MARA consistently improves retrieval relevance and answer quality over existing SOTA method. |
| title | MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering |
| topic | Information Retrieval Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2604.16313 |