From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Bhatia, Mehar, Ravi, Sahithya, Chinchure, Aditya, Hwang, Eunjeong, Shwartz, Vered |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SPIKE-RL: Video-LLMs meet Bayesian Surprise
por: Ravi, Sahithya, et al.
Publicado: (2025)
por: Ravi, Sahithya, et al.
Publicado: (2025)
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
por: Chinchure, Aditya, et al.
Publicado: (2025)
por: Chinchure, Aditya, et al.
Publicado: (2025)
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events
por: Chinchure, Aditya, et al.
Publicado: (2024)
por: Chinchure, Aditya, et al.
Publicado: (2024)
VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding
por: Yu, Haorui, et al.
Publicado: (2026)
por: Yu, Haorui, et al.
Publicado: (2026)
CulturalTeaming: AI-Assisted Interactive Red-Teaming for Challenging LLMs' (Lack of) Multicultural Knowledge
por: Chiu, Yu Ying, et al.
Publicado: (2024)
por: Chiu, Yu Ying, et al.
Publicado: (2024)
Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models
por: Jiang, Yifan, et al.
Publicado: (2026)
por: Jiang, Yifan, et al.
Publicado: (2026)
CROPE: Evaluating In-Context Adaptation of Vision and Language Models to Culture-Specific Concepts
por: Nikandrou, Malvina, et al.
Publicado: (2024)
por: Nikandrou, Malvina, et al.
Publicado: (2024)
Leveraging Vision-Language Pre-training for Human Activity Recognition in Still Images
por: Mahanta, Cristina, et al.
Publicado: (2025)
por: Mahanta, Cristina, et al.
Publicado: (2025)
Vision-Language Models Do Not Understand Negation
por: Alhamoud, Kumail, et al.
Publicado: (2025)
por: Alhamoud, Kumail, et al.
Publicado: (2025)
CulturalFrames: Assessing Cultural Expectation Alignment in Text-to-Image Models and Evaluation Metrics
por: Nayak, Shravan, et al.
Publicado: (2025)
por: Nayak, Shravan, et al.
Publicado: (2025)
Do Vision-Language Models Really Understand Visual Language?
por: Hou, Yifan, et al.
Publicado: (2024)
por: Hou, Yifan, et al.
Publicado: (2024)
Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly
por: Chetan, Aditya, et al.
Publicado: (2026)
por: Chetan, Aditya, et al.
Publicado: (2026)
Do Vision-Language Models Understand Compound Nouns?
por: Kumar, Sonal, et al.
Publicado: (2024)
por: Kumar, Sonal, et al.
Publicado: (2024)
Do Vision-Language Models Understand Visual Persuasiveness?
por: Park, Gyuwon
Publicado: (2025)
por: Park, Gyuwon
Publicado: (2025)
Intriguing Properties of Large Language and Vision Models
por: Lee, Young-Jun, et al.
Publicado: (2024)
por: Lee, Young-Jun, et al.
Publicado: (2024)
If CLIP Could Talk: Understanding Vision-Language Model Representations Through Their Preferred Concept Descriptions
por: Esfandiarpoor, Reza, et al.
Publicado: (2024)
por: Esfandiarpoor, Reza, et al.
Publicado: (2024)
CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation
por: Wang, Yuxuan, et al.
Publicado: (2024)
por: Wang, Yuxuan, et al.
Publicado: (2024)
Evaluating Vision-Language Models as Evaluators in Path Planning
por: Aghzal, Mohamed, et al.
Publicado: (2024)
por: Aghzal, Mohamed, et al.
Publicado: (2024)
PUMGPT: A Large Vision-Language Model for Product Understanding
por: Xue, Wei, et al.
Publicado: (2023)
por: Xue, Wei, et al.
Publicado: (2023)
Toward Interactive Regional Understanding in Vision-Large Language Models
por: Lee, Jungbeom, et al.
Publicado: (2024)
por: Lee, Jungbeom, et al.
Publicado: (2024)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
por: Wang, Xinyu, et al.
Publicado: (2025)
por: Wang, Xinyu, et al.
Publicado: (2025)
CartoMapQA: A Fundamental Benchmark Dataset Evaluating Vision-Language Models on Cartographic Map Understanding
por: Ung, Huy Quang, et al.
Publicado: (2025)
por: Ung, Huy Quang, et al.
Publicado: (2025)
BiasConnect: Investigating Bias Interactions in Text-to-Image Models
por: Shukla, Pushkar, et al.
Publicado: (2025)
por: Shukla, Pushkar, et al.
Publicado: (2025)
Evaluating Vision-Language Models for Emotion Recognition
por: Bhattacharyya, Sree, et al.
Publicado: (2025)
por: Bhattacharyya, Sree, et al.
Publicado: (2025)
Evaluation of Cultural Competence of Vision-Language Models
por: Yadav, Srishti, et al.
Publicado: (2025)
por: Yadav, Srishti, et al.
Publicado: (2025)
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models
por: Kim, Jeonghwan, et al.
Publicado: (2024)
por: Kim, Jeonghwan, et al.
Publicado: (2024)
Can Vision-Language Models Evaluate Handwritten Math?
por: Nath, Oikantik, et al.
Publicado: (2025)
por: Nath, Oikantik, et al.
Publicado: (2025)
Instruction-Following Evaluation of Large Vision-Language Models
por: Shiono, Daiki, et al.
Publicado: (2025)
por: Shiono, Daiki, et al.
Publicado: (2025)
ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models
por: Ruan, Chenxi, et al.
Publicado: (2026)
por: Ruan, Chenxi, et al.
Publicado: (2026)
An Explainable Biomedical Foundation Model via Large-Scale Concept-Enhanced Vision-Language Pre-training
por: Nie, Yuxiang, et al.
Publicado: (2025)
por: Nie, Yuxiang, et al.
Publicado: (2025)
Light Up the Shadows: Enhance Long-Tailed Entity Grounding with Concept-Guided Vision-Language Models
por: Zhang, Yikai, et al.
Publicado: (2024)
por: Zhang, Yikai, et al.
Publicado: (2024)
GeoCoder: Solving Geometry Problems by Generating Modular Code through Vision-Language Models
por: Sharma, Aditya, et al.
Publicado: (2024)
por: Sharma, Aditya, et al.
Publicado: (2024)
ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital Environments
por: Ray, Sourjyadip, et al.
Publicado: (2024)
por: Ray, Sourjyadip, et al.
Publicado: (2024)
Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models
por: Lu, Jiaying, et al.
Publicado: (2023)
por: Lu, Jiaying, et al.
Publicado: (2023)
Understanding Museum Exhibits using Vision-Language Reasoning
por: Balauca, Ada-Astrid, et al.
Publicado: (2024)
por: Balauca, Ada-Astrid, et al.
Publicado: (2024)
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
por: Verma, Arnav, et al.
Publicado: (2025)
por: Verma, Arnav, et al.
Publicado: (2025)
Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast Asia
por: Cahyawijaya, Samuel, et al.
Publicado: (2025)
por: Cahyawijaya, Samuel, et al.
Publicado: (2025)
HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning
por: Wei, Yanbin, et al.
Publicado: (2026)
por: Wei, Yanbin, et al.
Publicado: (2026)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
por: Zhu, Yingjie, et al.
Publicado: (2024)
por: Zhu, Yingjie, et al.
Publicado: (2024)
TIBET: Identifying and Evaluating Biases in Text-to-Image Generative Models
por: Chinchure, Aditya, et al.
Publicado: (2023)
por: Chinchure, Aditya, et al.
Publicado: (2023)
Ejemplares similares
-
SPIKE-RL: Video-LLMs meet Bayesian Surprise
por: Ravi, Sahithya, et al.
Publicado: (2025) -
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
por: Chinchure, Aditya, et al.
Publicado: (2025) -
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events
por: Chinchure, Aditya, et al.
Publicado: (2024) -
VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding
por: Yu, Haorui, et al.
Publicado: (2026) -
CulturalTeaming: AI-Assisted Interactive Red-Teaming for Challenging LLMs' (Lack of) Multicultural Knowledge
por: Chiu, Yu Ying, et al.
Publicado: (2024)