VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Haorui, Yang, Diji, He, Hang, Zhang, Fengrui, Yi, Qiufeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Seeing Symbols, Missing Cultures: Probing Vision-Language Models' Reasoning on Fire Imagery and Cultural Meaning
von: Yu, Haorui, et al.
Veröffentlicht: (2025)
von: Yu, Haorui, et al.
Veröffentlicht: (2025)
Cross-Cultural Expert-Level Art Critique Evaluation with Vision-Language Models
von: Yu, Haorui, et al.
Veröffentlicht: (2026)
von: Yu, Haorui, et al.
Veröffentlicht: (2026)
Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese
von: Inoue, Yuichi, et al.
Veröffentlicht: (2024)
von: Inoue, Yuichi, et al.
Veröffentlicht: (2024)
Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation
von: Zhou, Li, et al.
Veröffentlicht: (2025)
von: Zhou, Li, et al.
Veröffentlicht: (2025)
From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models
von: Bhatia, Mehar, et al.
Veröffentlicht: (2024)
von: Bhatia, Mehar, et al.
Veröffentlicht: (2024)
Benchmarking Vision Language Models for Cultural Understanding
von: Nayak, Shravan, et al.
Veröffentlicht: (2024)
von: Nayak, Shravan, et al.
Veröffentlicht: (2024)
CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
von: Wang, Zhaowei, et al.
Veröffentlicht: (2025)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models
von: Zhang, Junyi, et al.
Veröffentlicht: (2025)
von: Zhang, Junyi, et al.
Veröffentlicht: (2025)
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
von: Liu, Xuannan, et al.
Veröffentlicht: (2025)
von: Liu, Xuannan, et al.
Veröffentlicht: (2025)
VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
PM4Bench: Benchmarking Large Vision-Language Models with Parallel Multilingual Multi-Modal Multi-task Corpus
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)
MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
von: Wang, Shengkang, et al.
Veröffentlicht: (2024)
Evaluation of Cultural Competence of Vision-Language Models
von: Yadav, Srishti, et al.
Veröffentlicht: (2025)
von: Yadav, Srishti, et al.
Veröffentlicht: (2025)
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2026)
von: Faure, Gueter Josmy, et al.
Veröffentlicht: (2026)
HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
von: Li, Baiqi, et al.
Veröffentlicht: (2024)
von: Li, Baiqi, et al.
Veröffentlicht: (2024)
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
von: Li, Dingming, et al.
Veröffentlicht: (2025)
von: Li, Dingming, et al.
Veröffentlicht: (2025)
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling
von: Li, Siqi, et al.
Veröffentlicht: (2025)
von: Li, Siqi, et al.
Veröffentlicht: (2025)
CartoMapQA: A Fundamental Benchmark Dataset Evaluating Vision-Language Models on Cartographic Map Understanding
von: Ung, Huy Quang, et al.
Veröffentlicht: (2025)
von: Ung, Huy Quang, et al.
Veröffentlicht: (2025)
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
von: Li, Jiaang, et al.
Veröffentlicht: (2025)
von: Li, Jiaang, et al.
Veröffentlicht: (2025)
ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models
von: Ruan, Chenxi, et al.
Veröffentlicht: (2026)
von: Ruan, Chenxi, et al.
Veröffentlicht: (2026)
CDH-Bench: A Commonsense-Driven Hallucination Benchmark for Evaluating Visual Fidelity in Vision-Language Models
von: Chen, Kesheng, et al.
Veröffentlicht: (2026)
von: Chen, Kesheng, et al.
Veröffentlicht: (2026)
MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models
von: Zhou, Keyan, et al.
Veröffentlicht: (2025)
von: Zhou, Keyan, et al.
Veröffentlicht: (2025)
HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection
von: Gao, Siyuan, et al.
Veröffentlicht: (2025)
von: Gao, Siyuan, et al.
Veröffentlicht: (2025)
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
von: Kundu, Souvik, et al.
Veröffentlicht: (2025)
von: Kundu, Souvik, et al.
Veröffentlicht: (2025)
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
von: Chow, Wei, et al.
Veröffentlicht: (2025)
von: Chow, Wei, et al.
Veröffentlicht: (2025)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly
von: Chetan, Aditya, et al.
Veröffentlicht: (2026)
von: Chetan, Aditya, et al.
Veröffentlicht: (2026)
TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
von: He, Xingwei, et al.
Veröffentlicht: (2024)
von: He, Xingwei, et al.
Veröffentlicht: (2024)
CROPE: Evaluating In-Context Adaptation of Vision and Language Models to Culture-Specific Concepts
von: Nikandrou, Malvina, et al.
Veröffentlicht: (2024)
von: Nikandrou, Malvina, et al.
Veröffentlicht: (2024)
AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
von: Huang, Yipo, et al.
Veröffentlicht: (2024)
EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
von: Schäfer, Finn Rasmus, et al.
Veröffentlicht: (2026)
von: Schäfer, Finn Rasmus, et al.
Veröffentlicht: (2026)
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
von: Lim, Hyeonseok, et al.
Veröffentlicht: (2024)
von: Lim, Hyeonseok, et al.
Veröffentlicht: (2024)
SAP-Bench: Benchmarking Multimodal Large Language Models in Surgical Action Planning
von: Xu, Mengya, et al.
Veröffentlicht: (2025)
von: Xu, Mengya, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Seeing Symbols, Missing Cultures: Probing Vision-Language Models' Reasoning on Fire Imagery and Cultural Meaning
von: Yu, Haorui, et al.
Veröffentlicht: (2025) -
Cross-Cultural Expert-Level Art Critique Evaluation with Vision-Language Models
von: Yu, Haorui, et al.
Veröffentlicht: (2026) -
Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese
von: Inoue, Yuichi, et al.
Veröffentlicht: (2024) -
Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation
von: Zhou, Li, et al.
Veröffentlicht: (2025) -
From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models
von: Bhatia, Mehar, et al.
Veröffentlicht: (2024)