RISE-Video: Can Video Generators Decode Implicit World Rules?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Mingxin, Ma, Shuran, Meng, Shibei, Zhao, Xiangyu, Zhang, Zicheng, Zhang, Shaofeng, Zhong, Zhihang, Chen, Peixian, Cao, Haoyu, Sun, Xing, Duan, Haodong, Yang, Xue
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914308524343296
author Liu, Mingxin
Ma, Shuran
Meng, Shibei
Zhao, Xiangyu
Zhang, Zicheng
Zhang, Shaofeng
Zhong, Zhihang
Chen, Peixian
Cao, Haoyu
Sun, Xing
Duan, Haodong
Yang, Xue
author_facet Liu, Mingxin
Ma, Shuran
Meng, Shibei
Zhao, Xiangyu
Zhang, Zicheng
Zhang, Shaofeng
Zhong, Zhihang
Chen, Peixian
Cao, Haoyu
Sun, Xing
Duan, Haodong
Yang, Xue
contents While generative video models have achieved remarkable visual fidelity, their capacity to internalize and reason over implicit world rules remains a critical yet under-explored frontier. To bridge this gap, we present RISE-Video, a pioneering reasoning-oriented benchmark for Text-Image-to-Video (TI2V) synthesis that shifts the evaluative focus from surface-level aesthetics to deep cognitive reasoning. RISE-Video comprises 467 meticulously human-annotated samples spanning eight rigorous categories, providing a structured testbed for probing model intelligence across diverse dimensions, ranging from commonsense and spatial dynamics to specialized subject domains. Our framework introduces a multi-dimensional evaluation protocol consisting of four metrics: \textit{Reasoning Alignment}, \textit{Temporal Consistency}, \textit{Physical Rationality}, and \textit{Visual Quality}. To further support scalable evaluation, we propose an automated pipeline leveraging Large Multimodal Models (LMMs) to emulate human-centric assessment. Extensive experiments on 11 state-of-the-art TI2V models reveal pervasive deficiencies in simulating complex scenarios under implicit constraints, offering critical insights for the advancement of future world-simulating generative models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_05986
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RISE-Video: Can Video Generators Decode Implicit World Rules?
Liu, Mingxin
Ma, Shuran
Meng, Shibei
Zhao, Xiangyu
Zhang, Zicheng
Zhang, Shaofeng
Zhong, Zhihang
Chen, Peixian
Cao, Haoyu
Sun, Xing
Duan, Haodong
Yang, Xue
Computer Vision and Pattern Recognition
Artificial Intelligence
While generative video models have achieved remarkable visual fidelity, their capacity to internalize and reason over implicit world rules remains a critical yet under-explored frontier. To bridge this gap, we present RISE-Video, a pioneering reasoning-oriented benchmark for Text-Image-to-Video (TI2V) synthesis that shifts the evaluative focus from surface-level aesthetics to deep cognitive reasoning. RISE-Video comprises 467 meticulously human-annotated samples spanning eight rigorous categories, providing a structured testbed for probing model intelligence across diverse dimensions, ranging from commonsense and spatial dynamics to specialized subject domains. Our framework introduces a multi-dimensional evaluation protocol consisting of four metrics: \textit{Reasoning Alignment}, \textit{Temporal Consistency}, \textit{Physical Rationality}, and \textit{Visual Quality}. To further support scalable evaluation, we propose an automated pipeline leveraging Large Multimodal Models (LMMs) to emulate human-centric assessment. Extensive experiments on 11 state-of-the-art TI2V models reveal pervasive deficiencies in simulating complex scenarios under implicit constraints, offering critical insights for the advancement of future world-simulating generative models.
title RISE-Video: Can Video Generators Decode Implicit World Rules?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2602.05986