MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Darshan, Nagrani, Arsha, Manikantan, Kawshik, Singh, Harman, Tewari, Dinesh, Weyand, Tobias, Schmid, Cordelia, Angelova, Anelia, Dave, Shachi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MINERVA: Evaluating Complex Video Reasoning
by: Nagrani, Arsha, et al.
Published: (2025)
by: Nagrani, Arsha, et al.
Published: (2025)
CAViAR: Critic-Augmented Video Agentic Reasoning
by: Menon, Sachit, et al.
Published: (2025)
by: Menon, Sachit, et al.
Published: (2025)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
by: Min, Juhong, et al.
Published: (2024)
by: Min, Juhong, et al.
Published: (2024)
Neptune: The Long Orbit to Benchmarking Long Video Understanding
by: Nagrani, Arsha, et al.
Published: (2024)
by: Nagrani, Arsha, et al.
Published: (2024)
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
by: Nagrani, Arsha, et al.
Published: (2026)
by: Nagrani, Arsha, et al.
Published: (2026)
Streaming Dense Video Captioning
by: Zhou, Xingyi, et al.
Published: (2024)
by: Zhou, Xingyi, et al.
Published: (2024)
VoCap: Video Object Captioning and Segmentation from Any Prompt
by: Uijlings, Jasper, et al.
Published: (2025)
by: Uijlings, Jasper, et al.
Published: (2025)
SeeGULL Multilingual: a Dataset of Geo-Culturally Situated Stereotypes
by: Bhutani, Mukul, et al.
Published: (2024)
by: Bhutani, Mukul, et al.
Published: (2024)
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
by: Shvetsova, Nina, et al.
Published: (2025)
by: Shvetsova, Nina, et al.
Published: (2025)
VicTR: Video-conditioned Text Representations for Activity Recognition
by: Kahatapitiya, Kumara, et al.
Published: (2023)
by: Kahatapitiya, Kumara, et al.
Published: (2023)
Unlocking Multi-Spectral Data for Multi-Modal Models with Guided Inputs and Chain-of-Thought Reasoning
by: Kim, Dahun, et al.
Published: (2026)
by: Kim, Dahun, et al.
Published: (2026)
VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
by: Kim, Dahun, et al.
Published: (2025)
by: Kim, Dahun, et al.
Published: (2025)
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages
by: Singh, Harman, et al.
Published: (2024)
by: Singh, Harman, et al.
Published: (2024)
Time-Scaling State-Space Models for Dense Video Captioning
by: Piergiovanni, AJ, et al.
Published: (2025)
by: Piergiovanni, AJ, et al.
Published: (2025)
Region-centric Image-Language Pretraining for Open-Vocabulary Detection
by: Kim, Dahun, et al.
Published: (2023)
by: Kim, Dahun, et al.
Published: (2023)
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning
by: Piergiovanni, AJ, et al.
Published: (2024)
by: Piergiovanni, AJ, et al.
Published: (2024)
IdentifyMe: A Challenging Long-Context Mention Resolution Benchmark for LLMs
by: Manikantan, Kawshik, et al.
Published: (2024)
by: Manikantan, Kawshik, et al.
Published: (2024)
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
by: Ventura, Lucas, et al.
Published: (2025)
by: Ventura, Lucas, et al.
Published: (2025)
VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment
by: Saravanan, Darshana, et al.
Published: (2024)
by: Saravanan, Darshana, et al.
Published: (2024)
A Culturally-diverse Multilingual Multimodal Video Benchmark & Model
by: Shafique, Bhuiyan Sanjid, et al.
Published: (2025)
by: Shafique, Bhuiyan Sanjid, et al.
Published: (2025)
Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment
by: Kim, Dahun, et al.
Published: (2025)
by: Kim, Dahun, et al.
Published: (2025)
Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2024)
by: Kazakos, Evangelos, et al.
Published: (2024)
Beyond Aesthetics: Cultural Competence in Text-to-Image Models
by: Kannen, Nithish, et al.
Published: (2024)
by: Kannen, Nithish, et al.
Published: (2024)
Large-scale Pre-training for Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2025)
by: Kazakos, Evangelos, et al.
Published: (2025)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
by: Bousselham, Walid, et al.
Published: (2025)
by: Bousselham, Walid, et al.
Published: (2025)
AutoAD III: The Prequel -- Back to the Pixels
by: Han, Tengda, et al.
Published: (2024)
by: Han, Tengda, et al.
Published: (2024)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
by: Garcia, Ricardo, et al.
Published: (2024)
by: Garcia, Ricardo, et al.
Published: (2024)
SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels
by: Singh, Darshan, et al.
Published: (2024)
by: Singh, Darshan, et al.
Published: (2024)
Mirasol3B: A Multimodal Autoregressive model for time-aligned and contextual modalities
by: Piergiovanni, AJ, et al.
Published: (2023)
by: Piergiovanni, AJ, et al.
Published: (2023)
BrickNet: Graph-Backed Generative Brick Assembly
by: Kulits, Peter, et al.
Published: (2026)
by: Kulits, Peter, et al.
Published: (2026)
Dense Video Object Captioning from Disjoint Supervision
by: Zhou, Xingyi, et al.
Published: (2023)
by: Zhou, Xingyi, et al.
Published: (2023)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
by: Wysoczańska, Monika, et al.
Published: (2025)
by: Wysoczańska, Monika, et al.
Published: (2025)
CoVR-2: Automatic Data Construction for Composed Video Retrieval
by: Ventura, Lucas, et al.
Published: (2023)
by: Ventura, Lucas, et al.
Published: (2023)
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
by: Xie, Junyu, et al.
Published: (2024)
by: Xie, Junyu, et al.
Published: (2024)
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
by: Pacaud, Paul, et al.
Published: (2025)
by: Pacaud, Paul, et al.
Published: (2025)
RAVEN: Rethinking Adversarial Video Generation with Efficient Tri-plane Networks
by: Ghosh, Partha, et al.
Published: (2024)
by: Ghosh, Partha, et al.
Published: (2024)
Cross-Domain Identity Representation for Skull to Face Matching with Benchmark DataSet
by: Prasad, Ravi Shankar, et al.
Published: (2025)
by: Prasad, Ravi Shankar, et al.
Published: (2025)
Zero-Shot Multi-Spectral Learning: Reimagining a Generalist Multimodal Gemini 2.5 Model for Remote Sensing Applications
by: Mallya, Ganesh, et al.
Published: (2025)
by: Mallya, Ganesh, et al.
Published: (2025)
Major Entity Identification: A Generalizable Alternative to Coreference Resolution
by: Manikantan, Kawshik, et al.
Published: (2024)
by: Manikantan, Kawshik, et al.
Published: (2024)
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
by: Faraz, Ali, et al.
Published: (2025)
by: Faraz, Ali, et al.
Published: (2025)
Similar Items
-
MINERVA: Evaluating Complex Video Reasoning
by: Nagrani, Arsha, et al.
Published: (2025) -
CAViAR: Critic-Augmented Video Agentic Reasoning
by: Menon, Sachit, et al.
Published: (2025) -
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
by: Min, Juhong, et al.
Published: (2024) -
Neptune: The Long Orbit to Benchmarking Long Video Understanding
by: Nagrani, Arsha, et al.
Published: (2024) -
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
by: Nagrani, Arsha, et al.
Published: (2026)