WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shang, Yu, Tang, Yinzhou, Ma, Yiding, Li, Zhuohang, Jin, Lei, Su, Weikang, Jin, Xin, Wang, Zhaolu, Wang, Ziyou, Zhang, Xin, Su, Haisheng, He, Weizhen, Wu, Wei, Duan, Haoyi, Wetzstein, Gordon, Liu, Xihui, Shah, Dhruv, Zhang, Zhaoxiang, Chen, Zhibo, Zhu, Jun, Tian, Yonghong, Chua, Tat-Seng, Zhu, Wenwu, Gao, Chen, Li, Yong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913140172652544
author Shang, Yu
Tang, Yinzhou
Ma, Yiding
Li, Zhuohang
Jin, Lei
Su, Weikang
Jin, Xin
Wang, Zhaolu
Wang, Ziyou
Zhang, Xin
Su, Haisheng
He, Weizhen
Wu, Wei
Duan, Haoyi
Wetzstein, Gordon
Liu, Xihui
Shah, Dhruv
Zhang, Zhaoxiang
Chen, Zhibo
Zhu, Jun
Tian, Yonghong
Chua, Tat-Seng
Zhu, Wenwu
Gao, Chen
Li, Yong
author_facet Shang, Yu
Tang, Yinzhou
Ma, Yiding
Li, Zhuohang
Jin, Lei
Su, Weikang
Jin, Xin
Wang, Zhaolu
Wang, Ziyou
Zhang, Xin
Su, Haisheng
He, Weizhen
Wu, Wei
Duan, Haoyi
Wetzstein, Gordon
Liu, Xihui
Shah, Dhruv
Zhang, Zhaoxiang
Chen, Zhibo
Zhu, Jun
Tian, Yonghong
Chua, Tat-Seng
Zhu, Wenwu
Gao, Chen
Li, Yong
contents World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about environmental dynamics. However, existing embodied world model benchmarks are still largely confined to vision-only prediction, offline embodied applications, and simulator-based evaluation, making them insufficient for assessing increasingly comprehensive world models. In this work, we introduce WorldArena 2.0, an expanded benchmark that systematically broadens embodied world model evaluation along three dimensions: modality, functionality, and platform. Along the modality dimension, WorldArena 2.0 extends evaluation from vision-only to visuotactile modalities, enabling assessment of multimodal perception and prediction. Along the functionality dimension, it extends beyond policy evaluation and planning to assess world models as interactive RL environments for policy optimization. Along the platform dimension, it moves beyond simulator-only evaluation to a diverse suite of simulated and real-world robotic settings across multiple embodiments. Under a standardized protocol, WorldArena 2.0 comprehensively evaluates perceptual quality, interactive utility, and cross-platform performance, providing a comprehensive testbed for tracking progress toward embodied world models. The benchmark is available at: https://world-arena.ai.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17912
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform
Shang, Yu
Tang, Yinzhou
Ma, Yiding
Li, Zhuohang
Jin, Lei
Su, Weikang
Jin, Xin
Wang, Zhaolu
Wang, Ziyou
Zhang, Xin
Su, Haisheng
He, Weizhen
Wu, Wei
Duan, Haoyi
Wetzstein, Gordon
Liu, Xihui
Shah, Dhruv
Zhang, Zhaoxiang
Chen, Zhibo
Zhu, Jun
Tian, Yonghong
Chua, Tat-Seng
Zhu, Wenwu
Gao, Chen
Li, Yong
Robotics
Computer Vision and Pattern Recognition
World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about environmental dynamics. However, existing embodied world model benchmarks are still largely confined to vision-only prediction, offline embodied applications, and simulator-based evaluation, making them insufficient for assessing increasingly comprehensive world models. In this work, we introduce WorldArena 2.0, an expanded benchmark that systematically broadens embodied world model evaluation along three dimensions: modality, functionality, and platform. Along the modality dimension, WorldArena 2.0 extends evaluation from vision-only to visuotactile modalities, enabling assessment of multimodal perception and prediction. Along the functionality dimension, it extends beyond policy evaluation and planning to assess world models as interactive RL environments for policy optimization. Along the platform dimension, it moves beyond simulator-only evaluation to a diverse suite of simulated and real-world robotic settings across multiple embodiments. Under a standardized protocol, WorldArena 2.0 comprehensively evaluates perceptual quality, interactive utility, and cross-platform performance, providing a comprehensive testbed for tracking progress toward embodied world models. The benchmark is available at: https://world-arena.ai.
title WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.17912