Enhancing Document VQA Models via Retrieval-Augmented Generation
Fuente:
arXiv
Saved in:
| Main Authors: | López, Eric, Llabrés, Artemis, Valveny, Ernest |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Image-text matching for large-scale book collections
by: Llabrés, Artemis, et al.
Published: (2024)
by: Llabrés, Artemis, et al.
Published: (2024)
ComicsPAP: understanding comic strips by picking the correct panel
by: Vivoli, Emanuele, et al.
Published: (2025)
by: Vivoli, Emanuele, et al.
Published: (2025)
Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism
by: Kang, Lei, et al.
Published: (2024)
by: Kang, Lei, et al.
Published: (2024)
CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books
by: Ortega, Marc Serra, et al.
Published: (2025)
by: Ortega, Marc Serra, et al.
Published: (2025)
Retrieval Augmented Verification for Zero-Shot Detection of Multimodal Disinformation
by: Dey, Arka Ujjal, et al.
Published: (2024)
by: Dey, Arka Ujjal, et al.
Published: (2024)
Multimodal Retrieval-Augmented Generation with Large Language Models for Medical VQA
by: Karim, A H M Rezaul, et al.
Published: (2025)
by: Karim, A H M Rezaul, et al.
Published: (2025)
Machine Unlearning for Document Classification
by: Kang, Lei, et al.
Published: (2024)
by: Kang, Lei, et al.
Published: (2024)
Multimodal Transformer for Comics Text-Cloze
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
Reading in the Dark: Low-light Scene Text Recognition
by: Fu, Xuanshuo, et al.
Published: (2026)
by: Fu, Xuanshuo, et al.
Published: (2026)
Preserving Privacy Without Compromising Accuracy: Machine Unlearning for Handwritten Text Recognition
by: Kang, Lei, et al.
Published: (2025)
by: Kang, Lei, et al.
Published: (2025)
mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQA
by: Yuan, Xu, et al.
Published: (2025)
by: Yuan, Xu, et al.
Published: (2025)
LayeredDoc: Domain Adaptive Document Restoration with a Layer Separation Approach
by: Pilligua, Maria, et al.
Published: (2024)
by: Pilligua, Maria, et al.
Published: (2024)
GRIF-DM: Generation of Rich Impression Fonts using Diffusion Models
by: Kang, Lei, et al.
Published: (2024)
by: Kang, Lei, et al.
Published: (2024)
FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA
by: Sarwar, Nobin
Published: (2025)
by: Sarwar, Nobin
Published: (2025)
Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry
by: Hou, Wenjun, et al.
Published: (2024)
by: Hou, Wenjun, et al.
Published: (2024)
RAGME: Retrieval Augmented Video Generation for Enhanced Motion Realism
by: Peruzzo, Elia, et al.
Published: (2025)
by: Peruzzo, Elia, et al.
Published: (2025)
NeurIPS 2023 Competition: Privacy Preserving Federated Learning Document VQA
by: Tobaben, Marlon, et al.
Published: (2024)
by: Tobaben, Marlon, et al.
Published: (2024)
RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
by: Li, Yinglu, et al.
Published: (2025)
by: Li, Yinglu, et al.
Published: (2025)
SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs
by: Su, Xin, et al.
Published: (2024)
by: Su, Xin, et al.
Published: (2024)
Enhancing Vietnamese VQA through Curriculum Learning on Raw and Augmented Text Representations
by: Nguyen, Khoi Anh, et al.
Published: (2025)
by: Nguyen, Khoi Anh, et al.
Published: (2025)
A Simple Data Augmentation Strategy for Text-in-Image Scientific VQA
by: Shoer, Belal, et al.
Published: (2025)
by: Shoer, Belal, et al.
Published: (2025)
Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation
by: Xu, Quanxing, et al.
Published: (2026)
by: Xu, Quanxing, et al.
Published: (2026)
Retrieval Augmented Recipe Generation
by: Liu, Guoshan, et al.
Published: (2024)
by: Liu, Guoshan, et al.
Published: (2024)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
MobileRAG: Enhancing Mobile Agent with Retrieval-Augmented Generation
by: Loo, Gowen, et al.
Published: (2025)
by: Loo, Gowen, et al.
Published: (2025)
Exploring OCR-augmented Generation for Bilingual VQA
by: Lee, JoonHo, et al.
Published: (2025)
by: Lee, JoonHo, et al.
Published: (2025)
Enhancing Image Quality Assessment Ability of LMMs via Retrieval-Augmented Generation
by: Fu, Kang, et al.
Published: (2026)
by: Fu, Kang, et al.
Published: (2026)
Retrieval Augmented Comic Image Generation
by: Shui, Yunhao, et al.
Published: (2025)
by: Shui, Yunhao, et al.
Published: (2025)
E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation
by: Xu, Zeyu, et al.
Published: (2025)
by: Xu, Zeyu, et al.
Published: (2025)
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
by: Ma, Dongsheng, et al.
Published: (2026)
by: Ma, Dongsheng, et al.
Published: (2026)
VQA-Diff: Exploiting VQA and Diffusion for Zero-Shot Image-to-3D Vehicle Asset Generation in Autonomous Driving
by: Liu, Yibo, et al.
Published: (2024)
by: Liu, Yibo, et al.
Published: (2024)
J-RAS: Enhancing Medical Image Segmentation via Retrieval-Augmented Joint Training
by: Ahmed, Salma J., et al.
Published: (2025)
by: Ahmed, Salma J., et al.
Published: (2025)
VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
by: Chen, Jian, et al.
Published: (2025)
by: Chen, Jian, et al.
Published: (2025)
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
by: Gao, Sensen, et al.
Published: (2025)
by: Gao, Sensen, et al.
Published: (2025)
Hybrid Retrieval-Augmented Generation for Robust Multilingual Document Question Answering
by: Mudet, Anthony, et al.
Published: (2025)
by: Mudet, Anthony, et al.
Published: (2025)
Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?
by: Shen, Wenxuan, et al.
Published: (2025)
by: Shen, Wenxuan, et al.
Published: (2025)
Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
by: Dong, Kuicai, et al.
Published: (2025)
by: Dong, Kuicai, et al.
Published: (2025)
ASTRA: Enhancing Multi-Subject Generation with Retrieval-Augmented Pose Guidance and Disentangled Position Embedding
by: Xia, Tianze, et al.
Published: (2026)
by: Xia, Tianze, et al.
Published: (2026)
SimpsonsVQA: Enhancing Inquiry-Based Learning with a Tailored Dataset
by: Huynh, Ngoc Dung, et al.
Published: (2024)
by: Huynh, Ngoc Dung, et al.
Published: (2024)
Similar Items
-
Image-text matching for large-scale book collections
by: Llabrés, Artemis, et al.
Published: (2024) -
ComicsPAP: understanding comic strips by picking the correct panel
by: Vivoli, Emanuele, et al.
Published: (2025) -
Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism
by: Kang, Lei, et al.
Published: (2024) -
CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books
by: Ortega, Marc Serra, et al.
Published: (2025) -
Retrieval Augmented Verification for Zero-Shot Detection of Multimodal Disinformation
by: Dey, Arka Ujjal, et al.
Published: (2024)