Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Zhenghao, Zhu, Xingsheng, Zhou, Tianshuo, Zhang, Xinyi, Yi, Xiaoyuan, Yan, Yukun, Yu, Ge, Sun, Maosong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911095026876416
author Liu, Zhenghao
Zhu, Xingsheng
Zhou, Tianshuo
Zhang, Xinyi
Yi, Xiaoyuan
Yan, Yukun
Yu, Ge
Sun, Maosong
author_facet Liu, Zhenghao
Zhu, Xingsheng
Zhou, Tianshuo
Zhang, Xinyi
Yi, Xiaoyuan
Yan, Yukun
Yu, Ge
Sun, Maosong
contents With the rapid advancement of Multi-modal Large Language Models (MLLMs), their capability in understanding both images and text has greatly improved. However, their potential for leveraging multi-modal contextual information in Retrieval-Augmented Generation (RAG) remains largely underexplored. To address this gap, this paper introduces Multi-Modal Retrieval-Augmented Generation (M$^2$RAG), a benchmark designed to evaluate the effectiveness of Multi-modal Large Language Models in leveraging knowledge from multi-modal retrieval documents. The benchmark comprises four tasks: image captioning, multi-modal question answering, multi-modal fact verification, and image reranking. All tasks are set in an open-domain setting, requiring RAG models to retrieve query-relevant information from a multi-modal document collection and use it as contextual input for RAG modeling. To enhance the context utilization capabilities of MLLMs, we also introduce Multi-Modal Retrieval-Augmented Instruction Tuning (MM-RAIT), an instruction tuning method that optimizes MLLMs within multi-modal contexts. Our experiments demonstrate the effectiveness of MM-RAIT by significantly improving the quality of responses generated by different RAG models, outperforming MiniCPM-V 2.6 and Qwen2-VL with 34% and 33% gains, respectively. All data and code are available at https://github.com/NEUIR/M2RAG.
format Preprint
id arxiv_https___arxiv_org_abs_2502_17297
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts
Liu, Zhenghao
Zhu, Xingsheng
Zhou, Tianshuo
Zhang, Xinyi
Yi, Xiaoyuan
Yan, Yukun
Yu, Ge
Sun, Maosong
Artificial Intelligence
With the rapid advancement of Multi-modal Large Language Models (MLLMs), their capability in understanding both images and text has greatly improved. However, their potential for leveraging multi-modal contextual information in Retrieval-Augmented Generation (RAG) remains largely underexplored. To address this gap, this paper introduces Multi-Modal Retrieval-Augmented Generation (M$^2$RAG), a benchmark designed to evaluate the effectiveness of Multi-modal Large Language Models in leveraging knowledge from multi-modal retrieval documents. The benchmark comprises four tasks: image captioning, multi-modal question answering, multi-modal fact verification, and image reranking. All tasks are set in an open-domain setting, requiring RAG models to retrieve query-relevant information from a multi-modal document collection and use it as contextual input for RAG modeling. To enhance the context utilization capabilities of MLLMs, we also introduce Multi-Modal Retrieval-Augmented Instruction Tuning (MM-RAIT), an instruction tuning method that optimizes MLLMs within multi-modal contexts. Our experiments demonstrate the effectiveness of MM-RAIT by significantly improving the quality of responses generated by different RAG models, outperforming MiniCPM-V 2.6 and Qwen2-VL with 34% and 33% gains, respectively. All data and code are available at https://github.com/NEUIR/M2RAG.
title Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts
topic Artificial Intelligence
url https://arxiv.org/abs/2502.17297