WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Keming, Cui, Yijing, Xue, Wenhan, Wang, Qijie, Luo, Xuan, Feng, Zhiyuan, Yang, Zuhao, Wang, Sudong, Jiang, Sicong, Zhu, Haowei, Wang, Zihan, Nie, Ping, Chen, Wenhu, Wang, Bin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
by: Wu, Keming, et al.
Published: (2025)
by: Wu, Keming, et al.
Published: (2025)
Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling
by: Wu, Keming, et al.
Published: (2026)
by: Wu, Keming, et al.
Published: (2026)
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction
by: Liang, Jiarong, et al.
Published: (2026)
by: Liang, Jiarong, et al.
Published: (2026)
SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World Knowledge
by: Wang, Andong, et al.
Published: (2024)
by: Wang, Andong, et al.
Published: (2024)
MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks
by: Chen, Jiacheng, et al.
Published: (2024)
by: Chen, Jiacheng, et al.
Published: (2024)
Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem
by: Wang, Yubo, et al.
Published: (2025)
by: Wang, Yubo, et al.
Published: (2025)
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
by: Yang, Zuhao, et al.
Published: (2026)
by: Yang, Zuhao, et al.
Published: (2026)
ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks
by: Sani, Samin Mahdizadeh, et al.
Published: (2026)
by: Sani, Samin Mahdizadeh, et al.
Published: (2026)
UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing
by: Wang, Dianyi, et al.
Published: (2026)
by: Wang, Dianyi, et al.
Published: (2026)
ACECODER: Acing Coder RL via Automated Test-Case Synthesis
by: Zeng, Huaye, et al.
Published: (2025)
by: Zeng, Huaye, et al.
Published: (2025)
Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
by: Wang, Sudong, et al.
Published: (2026)
by: Wang, Sudong, et al.
Published: (2026)
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
by: Liang, Jiarong, et al.
Published: (2026)
by: Liang, Jiarong, et al.
Published: (2026)
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
by: Gao, Zeyu, et al.
Published: (2025)
by: Gao, Zeyu, et al.
Published: (2025)
Frozen Policy Iteration: Computationally Efficient RL under Linear $Q^π$ Realizability for Deterministic Dynamics
by: Ke, Yijing, et al.
Published: (2026)
by: Ke, Yijing, et al.
Published: (2026)
OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
by: Zhang, Kaichen, et al.
Published: (2025)
by: Zhang, Kaichen, et al.
Published: (2025)
Accelerating Diffusion-based Super-Resolution with Dynamic Time-Spatial Sampling
by: Qin, Rui, et al.
Published: (2025)
by: Qin, Rui, et al.
Published: (2025)
Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports
by: Yao, Yang, et al.
Published: (2025)
by: Yao, Yang, et al.
Published: (2025)
SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
by: Xu, Peiran, et al.
Published: (2025)
by: Xu, Peiran, et al.
Published: (2025)
WorldCoder-Bench: Benchmarking Physically Grounded 3D World Synthesis
by: Lu, Shuo, et al.
Published: (2026)
by: Lu, Shuo, et al.
Published: (2026)
Fine-grained Testing for Autonomous Driving Software: a Study on Autoware with LLM-driven Unit Testing
by: Wang, Wenhan, et al.
Published: (2025)
by: Wang, Wenhan, et al.
Published: (2025)
EvolveCoder: Evolving Test Cases via Adversarial Verification for Code Reinforcement Learning
by: Ruan, Chi, et al.
Published: (2026)
by: Ruan, Chi, et al.
Published: (2026)
VEglue: Testing Visual Entailment Systems via Object-Aligned Joint Erasing
by: Chang, Zhiyuan, et al.
Published: (2024)
by: Chang, Zhiyuan, et al.
Published: (2024)
Graph World Models: Concepts, Taxonomy, and Future Directions
by: Liu, Jiawei, et al.
Published: (2026)
by: Liu, Jiawei, et al.
Published: (2026)
VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
by: Wang, Kangrui, et al.
Published: (2025)
by: Wang, Kangrui, et al.
Published: (2025)
RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
Agglomeration Economies and Intergovernmental Cooperation in Productive Disaster Management Expenditure
by: Sudong Kim
Published: (2026)
by: Sudong Kim
Published: (2026)
WorldSimBench: Towards Video Generation Models as World Simulators
by: Qin, Yiran, et al.
Published: (2024)
by: Qin, Yiran, et al.
Published: (2024)
WorldModelBench: Judging Video Generation Models As World Models
by: Li, Dacheng, et al.
Published: (2025)
by: Li, Dacheng, et al.
Published: (2025)
WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
EponaV2: Driving World Model with Comprehensive Future Reasoning
by: Xu, Jiawei, et al.
Published: (2026)
by: Xu, Jiawei, et al.
Published: (2026)
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
by: Wang, Haozhe, et al.
Published: (2026)
by: Wang, Haozhe, et al.
Published: (2026)
PixelWorld: How Far Are We from Perceiving Everything as Pixels?
by: Lyu, Zhiheng, et al.
Published: (2025)
by: Lyu, Zhiheng, et al.
Published: (2025)
Real-Time Trend Prediction via Continually-Aligned LLM Query Generation
by: Hui, Zijing, et al.
Published: (2026)
by: Hui, Zijing, et al.
Published: (2026)
Polyelectrolyte‐Mediated Modulation of Spatial Internal Stresses of Hydrogels for Complex 3D Actuators
by: Jinghua Duan, et al.
Published: (2024)
by: Jinghua Duan, et al.
Published: (2024)
From Theory to Application: Fine-Tuning Large EEG Model with Real-World Stress Data
by: Wang, Siwen, et al.
Published: (2025)
by: Wang, Siwen, et al.
Published: (2025)
AdInject: Real-World Black-Box Attacks on Web Agents via Advertising Delivery
by: Wang, Haowei, et al.
Published: (2025)
by: Wang, Haowei, et al.
Published: (2025)
Open-World Test-Time Training: Self-Training with Contrast Learning
by: Su, Houcheng, et al.
Published: (2024)
by: Su, Houcheng, et al.
Published: (2024)
Can Test-Time Scaling Improve World Foundation Model?
by: Cong, Wenyan, et al.
Published: (2025)
by: Cong, Wenyan, et al.
Published: (2025)
Bongard-OpenWorld: Few-Shot Reasoning for Free-form Visual Concepts in the Real World
by: Wu, Rujie, et al.
Published: (2023)
by: Wu, Rujie, et al.
Published: (2023)
Similar Items
-
EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
by: Wu, Keming, et al.
Published: (2025) -
Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling
by: Wu, Keming, et al.
Published: (2026) -
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
by: Yang, Zuhao, et al.
Published: (2025) -
VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction
by: Liang, Jiarong, et al.
Published: (2026) -
SOK-Bench: A Situated Video Reasoning Benchmark with Aligned Open-World Knowledge
by: Wang, Andong, et al.
Published: (2024)