MINERVA: Evaluating Complex Video Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nagrani, Arsha, Menon, Sachit, Iscen, Ahmet, Buch, Shyamal, Mehran, Ramin, Jha, Nilpa, Hauth, Anja, Zhu, Yukun, Vondrick, Carl, Sirotenko, Mikhail, Schmid, Cordelia, Weyand, Tobias
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909598374428672
author Nagrani, Arsha
Menon, Sachit
Iscen, Ahmet
Buch, Shyamal
Mehran, Ramin
Jha, Nilpa
Hauth, Anja
Zhu, Yukun
Vondrick, Carl
Sirotenko, Mikhail
Schmid, Cordelia
Weyand, Tobias
author_facet Nagrani, Arsha
Menon, Sachit
Iscen, Ahmet
Buch, Shyamal
Mehran, Ramin
Jha, Nilpa
Hauth, Anja
Zhu, Yukun
Vondrick, Carl
Sirotenko, Mikhail
Schmid, Cordelia
Weyand, Tobias
contents Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps. This makes it challenging to assess if models are truly able to combine perceptual and temporal information to reason about videos, or simply get the correct answer by chance or by exploiting linguistic biases. To remedy this, we provide a new video reasoning dataset called MINERVA for modern multimodal models. Each question in the dataset comes with 5 answer choices, as well as detailed, hand-crafted reasoning traces. Our dataset is multimodal, diverse in terms of video domain and length, and consists of complex multi-step questions. Extensive benchmarking shows that our dataset provides a challenge for frontier open-source and proprietary models. We perform fine-grained error analysis to identify common failure modes across various models, and create a taxonomy of reasoning errors. We use this to explore both human and LLM-as-a-judge methods for scoring video reasoning traces, and find that failure modes are primarily related to temporal localization, followed by visual perception errors, as opposed to logical or completeness errors. The dataset, along with questions, answer candidates and reasoning traces will be publicly available under https://github.com/google-deepmind/neptune?tab=readme-ov-file\#minerva.
format Preprint
id arxiv_https___arxiv_org_abs_2505_00681
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MINERVA: Evaluating Complex Video Reasoning
Nagrani, Arsha
Menon, Sachit
Iscen, Ahmet
Buch, Shyamal
Mehran, Ramin
Jha, Nilpa
Hauth, Anja
Zhu, Yukun
Vondrick, Carl
Sirotenko, Mikhail
Schmid, Cordelia
Weyand, Tobias
Machine Learning
Computer Vision and Pattern Recognition
Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps. This makes it challenging to assess if models are truly able to combine perceptual and temporal information to reason about videos, or simply get the correct answer by chance or by exploiting linguistic biases. To remedy this, we provide a new video reasoning dataset called MINERVA for modern multimodal models. Each question in the dataset comes with 5 answer choices, as well as detailed, hand-crafted reasoning traces. Our dataset is multimodal, diverse in terms of video domain and length, and consists of complex multi-step questions. Extensive benchmarking shows that our dataset provides a challenge for frontier open-source and proprietary models. We perform fine-grained error analysis to identify common failure modes across various models, and create a taxonomy of reasoning errors. We use this to explore both human and LLM-as-a-judge methods for scoring video reasoning traces, and find that failure modes are primarily related to temporal localization, followed by visual perception errors, as opposed to logical or completeness errors. The dataset, along with questions, answer candidates and reasoning traces will be publicly available under https://github.com/google-deepmind/neptune?tab=readme-ov-file\#minerva.
title MINERVA: Evaluating Complex Video Reasoning
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.00681