WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Keming, Cui, Yijing, Xue, Wenhan, Wang, Qijie, Luo, Xuan, Feng, Zhiyuan, Yang, Zuhao, Wang, Sudong, Jiang, Sicong, Zhu, Haowei, Wang, Zihan, Nie, Ping, Chen, Wenhu, Wang, Bin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916000758235136
author Wu, Keming
Cui, Yijing
Xue, Wenhan
Wang, Qijie
Luo, Xuan
Feng, Zhiyuan
Yang, Zuhao
Wang, Sudong
Jiang, Sicong
Zhu, Haowei
Wang, Zihan
Nie, Ping
Chen, Wenhu
Wang, Bin
author_facet Wu, Keming
Cui, Yijing
Xue, Wenhan
Wang, Qijie
Luo, Xuan
Feng, Zhiyuan
Yang, Zuhao
Wang, Sudong
Jiang, Sicong
Zhu, Haowei
Wang, Zihan
Nie, Ping
Chen, Wenhu
Wang, Bin
contents Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving into "world simulators." Yet the community still lacks a benchmark that directly tests whether a model can reason about how an observed world should evolve over time. We introduce WorldReasonBench, which reframes video generation evaluation as world-state prediction: given an initial state and an action, can a model generate a future video whose state evolution remains physically, socially, logically, and informationally consistent? WorldReasonBench contains 436 curated test cases with structured ground-truth QA annotations spanning four reasoning dimensions and 22 subcategories. We evaluate generated videos with a human-aligned two-part methodology: Process-aware Reasoning Verification uses structured QA and reasoning-phase diagnostics to detect temporal and causal failures, while Multi-dimensional Quality Assessment scores reasoning quality, temporal consistency, and visual aesthetics for ranking and reward modeling. We further introduce WorldRewardBench, a preference benchmark with approximately 6K expert-annotated pairs over 1.4K videos, supporting pair-wise and point-wise reward-model evaluation. Across modern video generators, our results expose a persistent gap between visual plausibility and world reasoning: videos can look convincing while failing dynamics, causality, or information preservation. We will release our benchmarks and evaluation toolkit to support community research on genuinely world-aware video generation at https://github.com/UniX-AI-Lab/WorldReasonBench/.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10434
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors
Wu, Keming
Cui, Yijing
Xue, Wenhan
Wang, Qijie
Luo, Xuan
Feng, Zhiyuan
Yang, Zuhao
Wang, Sudong
Jiang, Sicong
Zhu, Haowei
Wang, Zihan
Nie, Ping
Chen, Wenhu
Wang, Bin
Computer Vision and Pattern Recognition
Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving into "world simulators." Yet the community still lacks a benchmark that directly tests whether a model can reason about how an observed world should evolve over time. We introduce WorldReasonBench, which reframes video generation evaluation as world-state prediction: given an initial state and an action, can a model generate a future video whose state evolution remains physically, socially, logically, and informationally consistent? WorldReasonBench contains 436 curated test cases with structured ground-truth QA annotations spanning four reasoning dimensions and 22 subcategories. We evaluate generated videos with a human-aligned two-part methodology: Process-aware Reasoning Verification uses structured QA and reasoning-phase diagnostics to detect temporal and causal failures, while Multi-dimensional Quality Assessment scores reasoning quality, temporal consistency, and visual aesthetics for ranking and reward modeling. We further introduce WorldRewardBench, a preference benchmark with approximately 6K expert-annotated pairs over 1.4K videos, supporting pair-wise and point-wise reward-model evaluation. Across modern video generators, our results expose a persistent gap between visual plausibility and world reasoning: videos can look convincing while failing dynamics, causality, or information preservation. We will release our benchmarks and evaluation toolkit to support community research on genuinely world-aware video generation at https://github.com/UniX-AI-Lab/WorldReasonBench/.
title WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.10434