MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singh, Darshan, Nagrani, Arsha, Manikantan, Kawshik, Singh, Harman, Tewari, Dinesh, Weyand, Tobias, Schmid, Cordelia, Angelova, Anelia, Dave, Shachi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917387310202880
author Singh, Darshan
Nagrani, Arsha
Manikantan, Kawshik
Singh, Harman
Tewari, Dinesh
Weyand, Tobias
Schmid, Cordelia
Angelova, Anelia
Dave, Shachi
author_facet Singh, Darshan
Nagrani, Arsha
Manikantan, Kawshik
Singh, Harman
Tewari, Dinesh
Weyand, Tobias
Schmid, Cordelia
Angelova, Anelia
Dave, Shachi
contents Recent advancements in video models have shown tremendous progress, particularly in long video understanding. However, current benchmarks predominantly feature western-centric data and English as the dominant language, introducing significant biases in evaluation. To address this, we introduce MINERVA-Cultural, a challenging benchmark for multicultural and multilingual video reasoning. MINERVA-Cultural comprises high-quality, entirely human-generated annotations from diverse, region-specific cultural videos across 18 global locales. Unlike prior work that relies on automatic translations, MINERVA-Cultural provides complex questions, answers, and multi-step reasoning steps, all crafted in native languages. Making progress on MINERVA-Cultural requires a deeply situated understanding of visual cultural context. Furthermore, we leverage MINERVA-Cultural's reasoning traces to construct evidence-based graphs and propose a novel iterative strategy using these graphs to identify fine-grained errors in reasoning. Our evaluations reveal that SoTA Video-LLMs struggle significantly, performing substantially below human-level accuracy, with errors primarily stemming from the visual perception of cultural elements. MINERVA-Cultural will be publicly available under https://github.com/google-deepmind/neptune?tab=readme-ov-file\#minerva-cultural
format Preprint
id arxiv_https___arxiv_org_abs_2601_10649
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
Singh, Darshan
Nagrani, Arsha
Manikantan, Kawshik
Singh, Harman
Tewari, Dinesh
Weyand, Tobias
Schmid, Cordelia
Angelova, Anelia
Dave, Shachi
Computer Vision and Pattern Recognition
Recent advancements in video models have shown tremendous progress, particularly in long video understanding. However, current benchmarks predominantly feature western-centric data and English as the dominant language, introducing significant biases in evaluation. To address this, we introduce MINERVA-Cultural, a challenging benchmark for multicultural and multilingual video reasoning. MINERVA-Cultural comprises high-quality, entirely human-generated annotations from diverse, region-specific cultural videos across 18 global locales. Unlike prior work that relies on automatic translations, MINERVA-Cultural provides complex questions, answers, and multi-step reasoning steps, all crafted in native languages. Making progress on MINERVA-Cultural requires a deeply situated understanding of visual cultural context. Furthermore, we leverage MINERVA-Cultural's reasoning traces to construct evidence-based graphs and propose a novel iterative strategy using these graphs to identify fine-grained errors in reasoning. Our evaluations reveal that SoTA Video-LLMs struggle significantly, performing substantially below human-level accuracy, with errors primarily stemming from the visual perception of cultural elements. MINERVA-Cultural will be publicly available under https://github.com/google-deepmind/neptune?tab=readme-ov-file\#minerva-cultural
title MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.10649