MINERVA: Evaluating Complex Video Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Nagrani, Arsha, Menon, Sachit, Iscen, Ahmet, Buch, Shyamal, Mehran, Ramin, Jha, Nilpa, Hauth, Anja, Zhu, Yukun, Vondrick, Carl, Sirotenko, Mikhail, Schmid, Cordelia, Weyand, Tobias |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CAViAR: Critic-Augmented Video Agentic Reasoning
by: Menon, Sachit, et al.
Published: (2025)
by: Menon, Sachit, et al.
Published: (2025)
Neptune: The Long Orbit to Benchmarking Long Video Understanding
by: Nagrani, Arsha, et al.
Published: (2024)
by: Nagrani, Arsha, et al.
Published: (2024)
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
by: Nagrani, Arsha, et al.
Published: (2026)
by: Nagrani, Arsha, et al.
Published: (2026)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
by: Min, Juhong, et al.
Published: (2024)
by: Min, Juhong, et al.
Published: (2024)
MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
by: Singh, Darshan, et al.
Published: (2026)
by: Singh, Darshan, et al.
Published: (2026)
Streaming Dense Video Captioning
by: Zhou, Xingyi, et al.
Published: (2024)
by: Zhou, Xingyi, et al.
Published: (2024)
Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities
by: Menon, Sachit, et al.
Published: (2024)
by: Menon, Sachit, et al.
Published: (2024)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
by: Wysoczańska, Monika, et al.
Published: (2025)
by: Wysoczańska, Monika, et al.
Published: (2025)
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
by: Caron, Mathilde, et al.
Published: (2024)
by: Caron, Mathilde, et al.
Published: (2024)
Retrieval-Enhanced Contrastive Vision-Text Models
by: Iscen, Ahmet, et al.
Published: (2023)
by: Iscen, Ahmet, et al.
Published: (2023)
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
by: Caron, Mathilde, et al.
Published: (2024)
by: Caron, Mathilde, et al.
Published: (2024)
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
by: Jain, Gagan, et al.
Published: (2024)
by: Jain, Gagan, et al.
Published: (2024)
Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
by: Arnab, Anurag, et al.
Published: (2025)
by: Arnab, Anurag, et al.
Published: (2025)
VoCap: Video Object Captioning and Segmentation from Any Prompt
by: Uijlings, Jasper, et al.
Published: (2025)
by: Uijlings, Jasper, et al.
Published: (2025)
Memory-Modular Classification: Learning to Generalize with Memory Replacement
by: Kang, Dahyun, et al.
Published: (2025)
by: Kang, Dahyun, et al.
Published: (2025)
Self-perceived Transformational Leadership Decreases Employee Sick Leave, but Context Matters
by: Tobias Hauth
Published: (2023)
by: Tobias Hauth
Published: (2023)
Continual Learning in Vision-Language Models via Aligned Model Merging
by: Sokar, Ghada, et al.
Published: (2025)
by: Sokar, Ghada, et al.
Published: (2025)
VicTR: Video-conditioned Text Representations for Activity Recognition
by: Kahatapitiya, Kumara, et al.
Published: (2023)
by: Kahatapitiya, Kumara, et al.
Published: (2023)
SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code
by: Hu, Ziniu, et al.
Published: (2024)
by: Hu, Ziniu, et al.
Published: (2024)
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
by: Shvetsova, Nina, et al.
Published: (2025)
by: Shvetsova, Nina, et al.
Published: (2025)
AutoAD III: The Prequel -- Back to the Pixels
by: Han, Tengda, et al.
Published: (2024)
by: Han, Tengda, et al.
Published: (2024)
CRIT: Graph-Based Automatic Data Synthesis to Enhance Cross-Modal Multi-Hop Reasoning
by: Sung, Junyoung, et al.
Published: (2026)
by: Sung, Junyoung, et al.
Published: (2026)
Generating Illustrated Instructions
by: Menon, Sachit, et al.
Published: (2023)
by: Menon, Sachit, et al.
Published: (2023)
BrickNet: Graph-Backed Generative Brick Assembly
by: Kulits, Peter, et al.
Published: (2026)
by: Kulits, Peter, et al.
Published: (2026)
4D Gaussian Splatting as a Learned Dynamical System
by: Asiimwe, Arnold Caleb, et al.
Published: (2025)
by: Asiimwe, Arnold Caleb, et al.
Published: (2025)
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
by: Xie, Junyu, et al.
Published: (2024)
by: Xie, Junyu, et al.
Published: (2024)
Problembezeichnung und Problemerlebnis - Gedanken zum problematischen Selbstverständnis einer Übersetzungswissenschaft
by: Ismail Işcen
Published: (2005)
by: Ismail Işcen
Published: (2005)
SelfIE: Self-Interpretation of Large Language Model Embeddings
by: Chen, Haozhe, et al.
Published: (2024)
by: Chen, Haozhe, et al.
Published: (2024)
Few-Shot Design Optimization by Exploiting Auxiliary Information
by: Mani, Arjun, et al.
Published: (2026)
by: Mani, Arjun, et al.
Published: (2026)
Evolving Interpretable Visual Classifiers with Large Language Models
by: Chiquier, Mia, et al.
Published: (2024)
by: Chiquier, Mia, et al.
Published: (2024)
Sobre la realidad de la vida cotidiana de los jóvenes en poblaciones en el nuevo orden democrático: «ni tan protagonista ni tan víctima»
by: Michaela Weyand
Published: (1993)
by: Michaela Weyand
Published: (1993)
Video Summarization: Towards Entity-Aware Captions
by: Ayyubi, Hammad A., et al.
Published: (2023)
by: Ayyubi, Hammad A., et al.
Published: (2023)
Extending Video Masked Autoencoders to 128 frames
by: Gundavarapu, Nitesh Bharadwaj, et al.
Published: (2024)
by: Gundavarapu, Nitesh Bharadwaj, et al.
Published: (2024)
Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation
by: Xie, Junyu, et al.
Published: (2025)
by: Xie, Junyu, et al.
Published: (2025)
More than a Moment: Towards Coherent Sequences of Audio Descriptions
by: Khandelwal, Eshika, et al.
Published: (2025)
by: Khandelwal, Eshika, et al.
Published: (2025)
Participatory provenance as representational auditing for AI-mediated public consultation
by: Mahajan, Sachit
Published: (2026)
by: Mahajan, Sachit
Published: (2026)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
by: Chen, Shizhe, et al.
Published: (2026)
by: Chen, Shizhe, et al.
Published: (2026)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
by: Bousselham, Walid, et al.
Published: (2025)
by: Bousselham, Walid, et al.
Published: (2025)
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
by: Fiastre, Gabriel, et al.
Published: (2025)
by: Fiastre, Gabriel, et al.
Published: (2025)
Learning text-to-video retrieval from image captioning
by: Ventura, Lucas, et al.
Published: (2024)
by: Ventura, Lucas, et al.
Published: (2024)
Similar Items
-
CAViAR: Critic-Augmented Video Agentic Reasoning
by: Menon, Sachit, et al.
Published: (2025) -
Neptune: The Long Orbit to Benchmarking Long Video Understanding
by: Nagrani, Arsha, et al.
Published: (2024) -
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
by: Nagrani, Arsha, et al.
Published: (2026) -
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
by: Min, Juhong, et al.
Published: (2024) -
MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
by: Singh, Darshan, et al.
Published: (2026)