ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Compagnoni, Alberto, Morini, Marco, Sarto, Sara, Cocchi, Federico, Caffagni, Davide, Cornia, Marcella, Baraldi, Lorenzo, Cucchiara, Rita
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912989592944640
author Compagnoni, Alberto
Morini, Marco
Sarto, Sara
Cocchi, Federico
Caffagni, Davide
Cornia, Marcella
Baraldi, Lorenzo
Cucchiara, Rita
author_facet Compagnoni, Alberto
Morini, Marco
Sarto, Sara
Cocchi, Federico
Caffagni, Davide
Cornia, Marcella
Baraldi, Lorenzo
Cucchiara, Rita
contents Multimodal Large Language Models (MLLMs) have shown impressive capabilities in jointly understanding text, images, and videos, often evaluated via Visual Question Answering (VQA). However, even state-of-the-art MLLMs struggle with domain-specific or knowledge-intensive queries, where relevant information is underrepresented in pre-training data. Knowledge-based VQA (KB-VQA) addresses this by retrieving external documents to condition answer generation, but current retrieval-augmented approaches suffer from low precision, noisy passages, and limited reasoning. To address this, we propose ReAG, a novel Reasoning-Augmented Multimodal RAG approach that combines coarse- and fine-grained retrieval with a critic model that filters irrelevant passages, ensuring high-quality additional context. The model follows a multi-stage training strategy leveraging reinforcement learning to enhance reasoning over retrieved content, while supervised fine-tuning serves only as a cold start. Extensive experiments on Encyclopedic-VQA and InfoSeek demonstrate that ReAG significantly outperforms prior methods, improving answer accuracy and providing interpretable reasoning grounded in retrieved evidence.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22715
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
Compagnoni, Alberto
Morini, Marco
Sarto, Sara
Cocchi, Federico
Caffagni, Davide
Cornia, Marcella
Baraldi, Lorenzo
Cucchiara, Rita
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
Multimodal Large Language Models (MLLMs) have shown impressive capabilities in jointly understanding text, images, and videos, often evaluated via Visual Question Answering (VQA). However, even state-of-the-art MLLMs struggle with domain-specific or knowledge-intensive queries, where relevant information is underrepresented in pre-training data. Knowledge-based VQA (KB-VQA) addresses this by retrieving external documents to condition answer generation, but current retrieval-augmented approaches suffer from low precision, noisy passages, and limited reasoning. To address this, we propose ReAG, a novel Reasoning-Augmented Multimodal RAG approach that combines coarse- and fine-grained retrieval with a critic model that filters irrelevant passages, ensuring high-quality additional context. The model follows a multi-stage training strategy leveraging reinforcement learning to enhance reasoning over retrieved content, while supervised fine-tuning serves only as a cold start. Extensive experiments on Encyclopedic-VQA and InfoSeek demonstrate that ReAG significantly outperforms prior methods, improving answer accuracy and providing interpretable reasoning grounded in retrieved evidence.
title ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimedia
url https://arxiv.org/abs/2511.22715