VideoNorms: Benchmarking Cultural Awareness of Video Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Varimalla, Nikhil Reddy, Xu, Yunfei, Saakyan, Arkadiy, Wang, Meng Fan, Muresan, Smaranda
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908585265463296
author Varimalla, Nikhil Reddy
Xu, Yunfei
Saakyan, Arkadiy
Wang, Meng Fan
Muresan, Smaranda
author_facet Varimalla, Nikhil Reddy
Xu, Yunfei
Saakyan, Arkadiy
Wang, Meng Fan
Muresan, Smaranda
contents As Video Large Language Models (VideoLLMs) are deployed globally, they require understanding of and grounding in the relevant cultural background. To properly assess these models' cultural awareness, adequate benchmarks are needed. We introduce VideoNorms, a benchmark of over 1000 (video clip, norm) pairs from US and Chinese cultures annotated with socio-cultural norms grounded in speech act theory, norm adherence and violations labels, and verbal and non-verbal evidence. To build VideoNorms, we use a human-AI collaboration framework, where a teacher model using theoretically-grounded prompting provides candidate annotations and a set of trained human experts validate and correct the annotations. We benchmark a variety of open-weight VideoLLMs on the new dataset which highlight several common trends: 1) models performs worse on norm violation than adherence; 2) models perform worse w.r.t Chinese culture compared to the US culture; 3) models have more difficulty in providing non-verbal evidence compared to verbal for the norm adhere/violation label and struggle to identify the exact norm corresponding to a speech-act; and 4) unlike humans, models perform worse in formal, non-humorous contexts. Our findings emphasize the need for culturally-grounded video language model training - a gap our benchmark and framework begin to address.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08543
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VideoNorms: Benchmarking Cultural Awareness of Video Language Models
Varimalla, Nikhil Reddy
Xu, Yunfei
Saakyan, Arkadiy
Wang, Meng Fan
Muresan, Smaranda
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Computers and Society
As Video Large Language Models (VideoLLMs) are deployed globally, they require understanding of and grounding in the relevant cultural background. To properly assess these models' cultural awareness, adequate benchmarks are needed. We introduce VideoNorms, a benchmark of over 1000 (video clip, norm) pairs from US and Chinese cultures annotated with socio-cultural norms grounded in speech act theory, norm adherence and violations labels, and verbal and non-verbal evidence. To build VideoNorms, we use a human-AI collaboration framework, where a teacher model using theoretically-grounded prompting provides candidate annotations and a set of trained human experts validate and correct the annotations. We benchmark a variety of open-weight VideoLLMs on the new dataset which highlight several common trends: 1) models performs worse on norm violation than adherence; 2) models perform worse w.r.t Chinese culture compared to the US culture; 3) models have more difficulty in providing non-verbal evidence compared to verbal for the norm adhere/violation label and struggle to identify the exact norm corresponding to a speech-act; and 4) unlike humans, models perform worse in formal, non-humorous contexts. Our findings emphasize the need for culturally-grounded video language model training - a gap our benchmark and framework begin to address.
title VideoNorms: Benchmarking Cultural Awareness of Video Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Computers and Society
url https://arxiv.org/abs/2510.08543