RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jiaang, Yuan, Yifei, Li, Wenyan, Aliannejadi, Mohammad, Hershcovich, Daniel, Søgaard, Anders, Vulić, Ivan, Zhang, Wenxuan, Liang, Paul Pu, Deng, Yang, Belongie, Serge
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908834929311744
author Li, Jiaang
Yuan, Yifei
Li, Wenyan
Aliannejadi, Mohammad
Hershcovich, Daniel
Søgaard, Anders
Vulić, Ivan
Zhang, Wenxuan
Liang, Paul Pu
Deng, Yang
Belongie, Serge
author_facet Li, Jiaang
Yuan, Yifei
Li, Wenyan
Aliannejadi, Mohammad
Hershcovich, Daniel
Søgaard, Anders
Vulić, Ivan
Zhang, Wenxuan
Liang, Paul Pu
Deng, Yang
Belongie, Serge
contents As vision-language models (VLMs) become increasingly integrated into daily life, the need for accurate visual culture understanding is becoming critical. Yet, these models frequently fall short in interpreting cultural nuances effectively. Prior work has demonstrated the effectiveness of retrieval-augmented generation (RAG) in enhancing cultural understanding in text-only settings, while its application in multimodal scenarios remains underexplored. To bridge this gap, we introduce RAVENEA (Retrieval-Augmented Visual culturE uNdErstAnding), a new benchmark designed to advance visual culture understanding through retrieval, focusing on two tasks: culture-focused visual question answering (cVQA) and culture-informed image captioning (cIC). RAVENEA extends existing datasets by integrating over 11,396 unique Wikipedia documents curated and ranked by human annotators. Through the extensive evaluation on seven multimodal retrievers and fifteen VLMs, RAVENEA reveals some undiscovered findings: (i) In general, cultural grounding annotations can enhance multimodal retrieval and corresponding downstream tasks. (ii) VLMs, when augmented with culture-aware retrieval, generally outperform their non-augmented counterparts (by averaging +6% on cVQA and +11% on cIC). (iii) Performance of culture-aware retrieval augmented varies widely across countries. These findings highlight the limitations of current multimodal retrievers and VLMs, underscoring the need to enhance visual culture understanding within RAG systems. We believe RAVENEA offers a valuable resource for advancing research on retrieval-augmented visual culture understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14462
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
Li, Jiaang
Yuan, Yifei
Li, Wenyan
Aliannejadi, Mohammad
Hershcovich, Daniel
Søgaard, Anders
Vulić, Ivan
Zhang, Wenxuan
Liang, Paul Pu
Deng, Yang
Belongie, Serge
Computer Vision and Pattern Recognition
Computation and Language
As vision-language models (VLMs) become increasingly integrated into daily life, the need for accurate visual culture understanding is becoming critical. Yet, these models frequently fall short in interpreting cultural nuances effectively. Prior work has demonstrated the effectiveness of retrieval-augmented generation (RAG) in enhancing cultural understanding in text-only settings, while its application in multimodal scenarios remains underexplored. To bridge this gap, we introduce RAVENEA (Retrieval-Augmented Visual culturE uNdErstAnding), a new benchmark designed to advance visual culture understanding through retrieval, focusing on two tasks: culture-focused visual question answering (cVQA) and culture-informed image captioning (cIC). RAVENEA extends existing datasets by integrating over 11,396 unique Wikipedia documents curated and ranked by human annotators. Through the extensive evaluation on seven multimodal retrievers and fifteen VLMs, RAVENEA reveals some undiscovered findings: (i) In general, cultural grounding annotations can enhance multimodal retrieval and corresponding downstream tasks. (ii) VLMs, when augmented with culture-aware retrieval, generally outperform their non-augmented counterparts (by averaging +6% on cVQA and +11% on cIC). (iii) Performance of culture-aware retrieval augmented varies widely across countries. These findings highlight the limitations of current multimodal retrievers and VLMs, underscoring the need to enhance visual culture understanding within RAG systems. We believe RAVENEA offers a valuable resource for advancing research on retrieval-augmented visual culture understanding.
title RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2505.14462