A Survey of Multimodal Retrieval-Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mei, Lang, Mo, Siyu, Yang, Zhihan, Chen, Chong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913789179330560
author Mei, Lang
Mo, Siyu
Yang, Zhihan
Chen, Chong
author_facet Mei, Lang
Mo, Siyu
Yang, Zhihan
Chen, Chong
contents Multimodal Retrieval-Augmented Generation (MRAG) enhances large language models (LLMs) by integrating multimodal data (text, images, videos) into retrieval and generation processes, overcoming the limitations of text-only Retrieval-Augmented Generation (RAG). While RAG improves response accuracy by incorporating external textual knowledge, MRAG extends this framework to include multimodal retrieval and generation, leveraging contextual information from diverse data types. This approach reduces hallucinations and enhances question-answering systems by grounding responses in factual, multimodal knowledge. Recent studies show MRAG outperforms traditional RAG, especially in scenarios requiring both visual and textual understanding. This survey reviews MRAG's essential components, datasets, evaluation methods, and limitations, providing insights into its construction and improvement. It also identifies challenges and future research directions, highlighting MRAG's potential to revolutionize multimodal information retrieval and generation. By offering a comprehensive perspective, this work encourages further exploration into this promising paradigm.
format Preprint
id arxiv_https___arxiv_org_abs_2504_08748
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Survey of Multimodal Retrieval-Augmented Generation
Mei, Lang
Mo, Siyu
Yang, Zhihan
Chen, Chong
Information Retrieval
Artificial Intelligence
Computation and Language
Emerging Technologies
Machine Learning
Multimodal Retrieval-Augmented Generation (MRAG) enhances large language models (LLMs) by integrating multimodal data (text, images, videos) into retrieval and generation processes, overcoming the limitations of text-only Retrieval-Augmented Generation (RAG). While RAG improves response accuracy by incorporating external textual knowledge, MRAG extends this framework to include multimodal retrieval and generation, leveraging contextual information from diverse data types. This approach reduces hallucinations and enhances question-answering systems by grounding responses in factual, multimodal knowledge. Recent studies show MRAG outperforms traditional RAG, especially in scenarios requiring both visual and textual understanding. This survey reviews MRAG's essential components, datasets, evaluation methods, and limitations, providing insights into its construction and improvement. It also identifies challenges and future research directions, highlighting MRAG's potential to revolutionize multimodal information retrieval and generation. By offering a comprehensive perspective, this work encourages further exploration into this promising paradigm.
title A Survey of Multimodal Retrieval-Augmented Generation
topic Information Retrieval
Artificial Intelligence
Computation and Language
Emerging Technologies
Machine Learning
url https://arxiv.org/abs/2504.08748