MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917387310202880 |
|---|---|
| author | Singh, Darshan Nagrani, Arsha Manikantan, Kawshik Singh, Harman Tewari, Dinesh Weyand, Tobias Schmid, Cordelia Angelova, Anelia Dave, Shachi |
| author_facet | Singh, Darshan Nagrani, Arsha Manikantan, Kawshik Singh, Harman Tewari, Dinesh Weyand, Tobias Schmid, Cordelia Angelova, Anelia Dave, Shachi |
| contents | Recent advancements in video models have shown tremendous progress, particularly in long video understanding. However, current benchmarks predominantly feature western-centric data and English as the dominant language, introducing significant biases in evaluation. To address this, we introduce MINERVA-Cultural, a challenging benchmark for multicultural and multilingual video reasoning. MINERVA-Cultural comprises high-quality, entirely human-generated annotations from diverse, region-specific cultural videos across 18 global locales. Unlike prior work that relies on automatic translations, MINERVA-Cultural provides complex questions, answers, and multi-step reasoning steps, all crafted in native languages. Making progress on MINERVA-Cultural requires a deeply situated understanding of visual cultural context. Furthermore, we leverage MINERVA-Cultural's reasoning traces to construct evidence-based graphs and propose a novel iterative strategy using these graphs to identify fine-grained errors in reasoning. Our evaluations reveal that SoTA Video-LLMs struggle significantly, performing substantially below human-level accuracy, with errors primarily stemming from the visual perception of cultural elements. MINERVA-Cultural will be publicly available under https://github.com/google-deepmind/neptune?tab=readme-ov-file\#minerva-cultural |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_10649 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning Singh, Darshan Nagrani, Arsha Manikantan, Kawshik Singh, Harman Tewari, Dinesh Weyand, Tobias Schmid, Cordelia Angelova, Anelia Dave, Shachi Computer Vision and Pattern Recognition Recent advancements in video models have shown tremendous progress, particularly in long video understanding. However, current benchmarks predominantly feature western-centric data and English as the dominant language, introducing significant biases in evaluation. To address this, we introduce MINERVA-Cultural, a challenging benchmark for multicultural and multilingual video reasoning. MINERVA-Cultural comprises high-quality, entirely human-generated annotations from diverse, region-specific cultural videos across 18 global locales. Unlike prior work that relies on automatic translations, MINERVA-Cultural provides complex questions, answers, and multi-step reasoning steps, all crafted in native languages. Making progress on MINERVA-Cultural requires a deeply situated understanding of visual cultural context. Furthermore, we leverage MINERVA-Cultural's reasoning traces to construct evidence-based graphs and propose a novel iterative strategy using these graphs to identify fine-grained errors in reasoning. Our evaluations reveal that SoTA Video-LLMs struggle significantly, performing substantially below human-level accuracy, with errors primarily stemming from the visual perception of cultural elements. MINERVA-Cultural will be publicly available under https://github.com/google-deepmind/neptune?tab=readme-ov-file\#minerva-cultural |
| title | MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2601.10649 |