RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jiaang, Yuan, Yifei, Li, Wenyan, Aliannejadi, Mohammad, Hershcovich, Daniel, Søgaard, Anders, Vulić, Ivan, Zhang, Wenxuan, Liang, Paul Pu, Deng, Yang, Belongie, Serge |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Visual Culture Awareness in GPT-4V: A Comprehensive Probing
by: Cao, Yong, et al.
Published: (2024)
by: Cao, Yong, et al.
Published: (2024)
FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food Culture
by: Li, Wenyan, et al.
Published: (2024)
by: Li, Wenyan, et al.
Published: (2024)
Unlocking Markets: A Multilingual Benchmark to Cross-Market Question Answering
by: Yuan, Yifei, et al.
Published: (2024)
by: Yuan, Yifei, et al.
Published: (2024)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
by: Li, Wenyan, et al.
Published: (2024)
by: Li, Wenyan, et al.
Published: (2024)
Lost in Embeddings: Information Loss in Vision-Language Models
by: Li, Wenyan, et al.
Published: (2025)
by: Li, Wenyan, et al.
Published: (2025)
Does Instruction Tuning Make LLMs More Consistent?
by: Fierro, Constanza, et al.
Published: (2024)
by: Fierro, Constanza, et al.
Published: (2024)
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
by: Li, Chengzu, et al.
Published: (2026)
by: Li, Chengzu, et al.
Published: (2026)
What if Othello-Playing Language Models Could See?
by: Chen, Xinyi, et al.
Published: (2025)
by: Chen, Xinyi, et al.
Published: (2025)
Do Vision and Language Models Share Concepts? A Vector Space Alignment Study
by: Li, Jiaang, et al.
Published: (2023)
by: Li, Jiaang, et al.
Published: (2023)
Evaluation of Cultural Competence of Vision-Language Models
by: Yadav, Srishti, et al.
Published: (2025)
by: Yadav, Srishti, et al.
Published: (2025)
Vision-Language Models under Cultural and Inclusive Considerations
by: Karamolegkou, Antonia, et al.
Published: (2024)
by: Karamolegkou, Antonia, et al.
Published: (2024)
Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired Users
by: Karamolegkou, Antonia, et al.
Published: (2025)
by: Karamolegkou, Antonia, et al.
Published: (2025)
Query Understanding in LLM-based Conversational Information Seeking
by: Yuan, Yifei, et al.
Published: (2025)
by: Yuan, Yifei, et al.
Published: (2025)
Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation
by: Zhou, Li, et al.
Published: (2025)
by: Zhou, Li, et al.
Published: (2025)
Revisiting the Othello World Model Hypothesis
by: Yuan, Yifei, et al.
Published: (2025)
by: Yuan, Yifei, et al.
Published: (2025)
ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question Answering
by: Lupart, Simon, et al.
Published: (2025)
by: Lupart, Simon, et al.
Published: (2025)
Do LLMs Understand Wine Descriptors Across Cultures? A Benchmark for Cultural Adaptations of Wine Reviews
by: Zou, Chenye, et al.
Published: (2025)
by: Zou, Chenye, et al.
Published: (2025)
Video Understanding: From Geometry and Semantics to Unified Models
by: An, Zhaochong, et al.
Published: (2026)
by: An, Zhaochong, et al.
Published: (2026)
EvalCards: A Framework for Standardized Evaluation Reporting
by: Dhar, Ruchira, et al.
Published: (2025)
by: Dhar, Ruchira, et al.
Published: (2025)
Asking Multimodal Clarifying Questions in Mixed-Initiative Conversational Search
by: Yuan, Yifei, et al.
Published: (2024)
by: Yuan, Yifei, et al.
Published: (2024)
Align Documents to Questions: Question-Oriented Document Rewriting for Retrieval-Augmented Generation
by: Li, Jiaang, et al.
Published: (2026)
by: Li, Jiaang, et al.
Published: (2026)
HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals
by: Hu, Guimin, et al.
Published: (2025)
by: Hu, Guimin, et al.
Published: (2025)
MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training
by: Gordon, Lucia, et al.
Published: (2026)
by: Gordon, Lucia, et al.
Published: (2026)
Self-Augmented In-Context Learning for Unsupervised Word Translation
by: Li, Yaoyiran, et al.
Published: (2024)
by: Li, Yaoyiran, et al.
Published: (2024)
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
by: Li, Chengzu, et al.
Published: (2025)
by: Li, Chengzu, et al.
Published: (2025)
Understanding Retrieval-Augmented Task Adaptation for Vision-Language Models
by: Ming, Yifei, et al.
Published: (2024)
by: Ming, Yifei, et al.
Published: (2024)
Cultural Compass: Predicting Transfer Learning Success in Offensive Language Detection with Cultural Features
by: Zhou, Li, et al.
Published: (2023)
by: Zhou, Li, et al.
Published: (2023)
Better Language Models Exhibit Higher Visual Alignment
by: Ruthardt, Jona, et al.
Published: (2024)
by: Ruthardt, Jona, et al.
Published: (2024)
Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models
by: Yadav, Srishti, et al.
Published: (2025)
by: Yadav, Srishti, et al.
Published: (2025)
Bridging Cultural Nuances in Dialogue Agents through Cultural Value Surveys
by: Cao, Yong, et al.
Published: (2024)
by: Cao, Yong, et al.
Published: (2024)
Learning to Judge: LLMs Designing and Applying Evaluation Rubrics
by: Siro, Clemencia, et al.
Published: (2026)
by: Siro, Clemencia, et al.
Published: (2026)
Fleurs-SLU: A Massively Multilingual Benchmark for Spoken Language Understanding
by: Schmidt, Fabian David, et al.
Published: (2025)
by: Schmidt, Fabian David, et al.
Published: (2025)
DiSCo: LLM Knowledge Distillation for Efficient Sparse Retrieval in Conversational Search
by: Lupart, Simon, et al.
Published: (2024)
by: Lupart, Simon, et al.
Published: (2024)
DARE: Diverse Visual Question Answering with Robustness Evaluation
by: Sterz, Hannah, et al.
Published: (2024)
by: Sterz, Hannah, et al.
Published: (2024)
From Words to Worlds: Compositionality for Cognitive Architectures
by: Dhar, Ruchira, et al.
Published: (2024)
by: Dhar, Ruchira, et al.
Published: (2024)
Understanding Subword Compositionality of Large Language Models
by: Peng, Qiwei, et al.
Published: (2025)
by: Peng, Qiwei, et al.
Published: (2025)
IRLab@iKAT24: Learned Sparse Retrieval with Multi-aspect LLM Query Generation for Conversational Search
by: Lupart, Simon, et al.
Published: (2024)
by: Lupart, Simon, et al.
Published: (2024)
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
by: Yu, Hong-Tao, et al.
Published: (2025)
by: Yu, Hong-Tao, et al.
Published: (2025)
Multi-Turn Multi-Modal Question Clarification for Enhanced Conversational Understanding
by: Ramezan, Kimia, et al.
Published: (2025)
by: Ramezan, Kimia, et al.
Published: (2025)
Hypencoder Revisited: Reproducibility and Analysis of Non-Linear Scoring for First-Stage Retrieval
by: Eichholtz, Arne, et al.
Published: (2026)
by: Eichholtz, Arne, et al.
Published: (2026)
Similar Items
-
Exploring Visual Culture Awareness in GPT-4V: A Comprehensive Probing
by: Cao, Yong, et al.
Published: (2024) -
FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food Culture
by: Li, Wenyan, et al.
Published: (2024) -
Unlocking Markets: A Multilingual Benchmark to Cross-Market Question Answering
by: Yuan, Yifei, et al.
Published: (2024) -
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
by: Li, Wenyan, et al.
Published: (2024) -
Lost in Embeddings: Information Loss in Vision-Language Models
by: Li, Wenyan, et al.
Published: (2025)