World Reasoning Arena

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: PAN Team, Gao, Qiyue, Zhou, Kun, Xiang, Jiannan, Liu, Zihan, Yang, Dequan, Chen, Junrong, Ahmad, Arif, Zeng, Cong, Bannur, Ganesh, Huang, Xinqi, Liu, Zheqi, Gu, Yi, Yang, Yichi, Liu, Guangyi, Hu, Zhiting, Liu, Zhengzhong, Xing, Eric
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915893276049408
author PAN Team
Gao, Qiyue
Zhou, Kun
Xiang, Jiannan
Liu, Zihan
Yang, Dequan
Chen, Junrong
Ahmad, Arif
Zeng, Cong
Bannur, Ganesh
Huang, Xinqi
Liu, Zheqi
Gu, Yi
Yang, Yichi
Liu, Guangyi
Hu, Zhiting
Liu, Zhengzhong
Xing, Eric
author_facet PAN Team
Gao, Qiyue
Zhou, Kun
Xiang, Jiannan
Liu, Zihan
Yang, Dequan
Chen, Junrong
Ahmad, Arif
Zeng, Cong
Bannur, Ganesh
Huang, Xinqi
Liu, Zheqi
Gu, Yi
Yang, Yichi
Liu, Guangyi
Hu, Zhiting
Liu, Zhengzhong
Xing, Eric
contents World models (WMs) are intended to serve as internal simulators of the real world that enable agents to understand, anticipate, and act upon complex environments. Existing WM benchmarks remain narrowly focused on next-state prediction and visual fidelity, overlooking the richer simulation capabilities required for intelligent behavior. To address this gap, we introduce WR-Arena, a comprehensive benchmark for evaluating WMs along three fundamental dimensions of next world simulation: (i) Action Simulation Fidelity, the ability to interpret and follow semantically meaningful, multi-step instructions and generate diverse counterfactual rollouts; (ii) Long-horizon Forecast, the ability to sustain accurate, coherent, and physically plausible simulations across extended interactions; and (iii) Simulative Reasoning and Planning, the ability to support goal-directed reasoning by simulating, comparing, and selecting among alternative futures in both structured and open-ended environments. We build a task taxonomy and curate diverse datasets designed to probe these capabilities, moving beyond single-turn and perceptual evaluations. Through extensive experiments with state-of-the-art WMs, our results expose a substantial gap between current models and human-level hypothetical reasoning, and establish WR-Arena as both a diagnostic tool and a guideline for advancing next-generation world models capable of robust understanding, forecasting, and purposeful action. The code is available at https://github.com/MBZUAI-IFM/WR-Arena.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25887
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle World Reasoning Arena
PAN Team
Gao, Qiyue
Zhou, Kun
Xiang, Jiannan
Liu, Zihan
Yang, Dequan
Chen, Junrong
Ahmad, Arif
Zeng, Cong
Bannur, Ganesh
Huang, Xinqi
Liu, Zheqi
Gu, Yi
Yang, Yichi
Liu, Guangyi
Hu, Zhiting
Liu, Zhengzhong
Xing, Eric
Computer Vision and Pattern Recognition
World models (WMs) are intended to serve as internal simulators of the real world that enable agents to understand, anticipate, and act upon complex environments. Existing WM benchmarks remain narrowly focused on next-state prediction and visual fidelity, overlooking the richer simulation capabilities required for intelligent behavior. To address this gap, we introduce WR-Arena, a comprehensive benchmark for evaluating WMs along three fundamental dimensions of next world simulation: (i) Action Simulation Fidelity, the ability to interpret and follow semantically meaningful, multi-step instructions and generate diverse counterfactual rollouts; (ii) Long-horizon Forecast, the ability to sustain accurate, coherent, and physically plausible simulations across extended interactions; and (iii) Simulative Reasoning and Planning, the ability to support goal-directed reasoning by simulating, comparing, and selecting among alternative futures in both structured and open-ended environments. We build a task taxonomy and curate diverse datasets designed to probe these capabilities, moving beyond single-turn and perceptual evaluations. Through extensive experiments with state-of-the-art WMs, our results expose a substantial gap between current models and human-level hypothetical reasoning, and establish WR-Arena as both a diagnostic tool and a guideline for advancing next-generation world models capable of robust understanding, forecasting, and purposeful action. The code is available at https://github.com/MBZUAI-IFM/WR-Arena.
title World Reasoning Arena
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.25887