WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913140172652544 |
|---|---|
| author | Shang, Yu Tang, Yinzhou Ma, Yiding Li, Zhuohang Jin, Lei Su, Weikang Jin, Xin Wang, Zhaolu Wang, Ziyou Zhang, Xin Su, Haisheng He, Weizhen Wu, Wei Duan, Haoyi Wetzstein, Gordon Liu, Xihui Shah, Dhruv Zhang, Zhaoxiang Chen, Zhibo Zhu, Jun Tian, Yonghong Chua, Tat-Seng Zhu, Wenwu Gao, Chen Li, Yong |
| author_facet | Shang, Yu Tang, Yinzhou Ma, Yiding Li, Zhuohang Jin, Lei Su, Weikang Jin, Xin Wang, Zhaolu Wang, Ziyou Zhang, Xin Su, Haisheng He, Weizhen Wu, Wei Duan, Haoyi Wetzstein, Gordon Liu, Xihui Shah, Dhruv Zhang, Zhaoxiang Chen, Zhibo Zhu, Jun Tian, Yonghong Chua, Tat-Seng Zhu, Wenwu Gao, Chen Li, Yong |
| contents | World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about environmental dynamics. However, existing embodied world model benchmarks are still largely confined to vision-only prediction, offline embodied applications, and simulator-based evaluation, making them insufficient for assessing increasingly comprehensive world models. In this work, we introduce WorldArena 2.0, an expanded benchmark that systematically broadens embodied world model evaluation along three dimensions: modality, functionality, and platform. Along the modality dimension, WorldArena 2.0 extends evaluation from vision-only to visuotactile modalities, enabling assessment of multimodal perception and prediction. Along the functionality dimension, it extends beyond policy evaluation and planning to assess world models as interactive RL environments for policy optimization. Along the platform dimension, it moves beyond simulator-only evaluation to a diverse suite of simulated and real-world robotic settings across multiple embodiments. Under a standardized protocol, WorldArena 2.0 comprehensively evaluates perceptual quality, interactive utility, and cross-platform performance, providing a comprehensive testbed for tracking progress toward embodied world models. The benchmark is available at: https://world-arena.ai. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_17912 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform Shang, Yu Tang, Yinzhou Ma, Yiding Li, Zhuohang Jin, Lei Su, Weikang Jin, Xin Wang, Zhaolu Wang, Ziyou Zhang, Xin Su, Haisheng He, Weizhen Wu, Wei Duan, Haoyi Wetzstein, Gordon Liu, Xihui Shah, Dhruv Zhang, Zhaoxiang Chen, Zhibo Zhu, Jun Tian, Yonghong Chua, Tat-Seng Zhu, Wenwu Gao, Chen Li, Yong Robotics Computer Vision and Pattern Recognition World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about environmental dynamics. However, existing embodied world model benchmarks are still largely confined to vision-only prediction, offline embodied applications, and simulator-based evaluation, making them insufficient for assessing increasingly comprehensive world models. In this work, we introduce WorldArena 2.0, an expanded benchmark that systematically broadens embodied world model evaluation along three dimensions: modality, functionality, and platform. Along the modality dimension, WorldArena 2.0 extends evaluation from vision-only to visuotactile modalities, enabling assessment of multimodal perception and prediction. Along the functionality dimension, it extends beyond policy evaluation and planning to assess world models as interactive RL environments for policy optimization. Along the platform dimension, it moves beyond simulator-only evaluation to a diverse suite of simulated and real-world robotic settings across multiple embodiments. Under a standardized protocol, WorldArena 2.0 comprehensively evaluates perceptual quality, interactive utility, and cross-platform performance, providing a comprehensive testbed for tracking progress toward embodied world models. The benchmark is available at: https://world-arena.ai. |
| title | WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform |
| topic | Robotics Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.17912 |