V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866911277567180800 |
|---|---|
| author | Luo, Yang Zhao, Xuanlei Lin, Baijiong Zhu, Lingting Tang, Liyao Liu, Yuqi Chen, Ying-Cong Qian, Shengju Wang, Xin You, Yang |
| author_facet | Luo, Yang Zhao, Xuanlei Lin, Baijiong Zhu, Lingting Tang, Liyao Liu, Yuqi Chen, Ying-Cong Qian, Shengju Wang, Xin You, Yang |
| contents | Recent progress in generative video models, such as Veo-3, has shown surprising zero-shot reasoning abilities, creating a growing need for systematic and reliable evaluation. We introduce V-ReasonBench, a benchmark designed to assess video reasoning across four key dimensions: structured problem-solving, spatial cognition, pattern-based inference, and physical dynamics. The benchmark is built from both synthetic and real-world image sequences and provides a diverse set of answer-verifiable tasks that are reproducible, scalable, and unambiguous. Evaluations of six state-of-the-art video models reveal clear dimension-wise differences, with strong variation in structured, spatial, pattern-based, and physical reasoning. We further compare video models with strong image models, analyze common hallucination behaviors, and study how video duration affects Chain-of-Frames reasoning. Overall, V-ReasonBench offers a unified and reproducible framework for measuring video reasoning and aims to support the development of models with more reliable, human-aligned reasoning skills. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_16668 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models Luo, Yang Zhao, Xuanlei Lin, Baijiong Zhu, Lingting Tang, Liyao Liu, Yuqi Chen, Ying-Cong Qian, Shengju Wang, Xin You, Yang Computer Vision and Pattern Recognition Recent progress in generative video models, such as Veo-3, has shown surprising zero-shot reasoning abilities, creating a growing need for systematic and reliable evaluation. We introduce V-ReasonBench, a benchmark designed to assess video reasoning across four key dimensions: structured problem-solving, spatial cognition, pattern-based inference, and physical dynamics. The benchmark is built from both synthetic and real-world image sequences and provides a diverse set of answer-verifiable tasks that are reproducible, scalable, and unambiguous. Evaluations of six state-of-the-art video models reveal clear dimension-wise differences, with strong variation in structured, spatial, pattern-based, and physical reasoning. We further compare video models with strong image models, analyze common hallucination behaviors, and study how video duration affects Chain-of-Frames reasoning. Overall, V-ReasonBench offers a unified and reproducible framework for measuring video reasoning and aims to support the development of models with more reliable, human-aligned reasoning skills. |
| title | V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.16668 |