ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911553720156160 |
|---|---|
| author | Sun, Yu Cao, Meng Yang, Ping Xu, Rongtao Yan, Yunxiao Xu, Runze Ma, Liang Gan, Roy Zhai, Andy Chen, Qingxuan Xu, Zunnan Wang, Hao Yu, Jincheng Liang, Lucy Wang, Qian Laptev, Ivan Reid, Ian D Liang, Xiaodan |
| author_facet | Sun, Yu Cao, Meng Yang, Ping Xu, Rongtao Yan, Yunxiao Xu, Runze Ma, Liang Gan, Roy Zhai, Andy Chen, Qingxuan Xu, Zunnan Wang, Hao Yu, Jincheng Liang, Lucy Wang, Qian Laptev, Ivan Reid, Ian D Liang, Xiaodan |
| contents | Vision-Language-Action (VLA) models and world models have recently emerged as promising paradigms for general-purpose robotic intelligence, yet their progress is hindered by the lack of reliable evaluation protocols that reflect real-world deployment. Existing benchmarks are largely simulator-centric, which provide controllability but fail to capture the reality gap caused by perception noise, complex contact dynamics, hardware constraints, and system latency. Moreover, fragmented real-world evaluations across different robot platforms prevent fair and reproducible comparison. To address these challenges, we introduce ManipArena, a standardized evaluation framework designed to bridge simulation and real-world execution. ManipArena comprises 20 diverse tasks across 10,812 expert trajectories emphasizing reasoning-oriented manipulation tasks requiring semantic and spatial reasoning, supports multi-level generalization through controlled out-of-distribution settings, and incorporates long-horizon mobile manipulation beyond tabletop scenarios. The framework further provides rich sensory diagnostics, including low-level motor signals, and synchronized real-to-sim environments constructed via high-quality 3D scanning. Together, these features enable fair, realistic, and reproducible evaluation for both VLA and world model approaches, providing a scalable foundation for diagnosing and advancing embodied intelligence systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_28545 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation Sun, Yu Cao, Meng Yang, Ping Xu, Rongtao Yan, Yunxiao Xu, Runze Ma, Liang Gan, Roy Zhai, Andy Chen, Qingxuan Xu, Zunnan Wang, Hao Yu, Jincheng Liang, Lucy Wang, Qian Laptev, Ivan Reid, Ian D Liang, Xiaodan Robotics Computer Vision and Pattern Recognition Vision-Language-Action (VLA) models and world models have recently emerged as promising paradigms for general-purpose robotic intelligence, yet their progress is hindered by the lack of reliable evaluation protocols that reflect real-world deployment. Existing benchmarks are largely simulator-centric, which provide controllability but fail to capture the reality gap caused by perception noise, complex contact dynamics, hardware constraints, and system latency. Moreover, fragmented real-world evaluations across different robot platforms prevent fair and reproducible comparison. To address these challenges, we introduce ManipArena, a standardized evaluation framework designed to bridge simulation and real-world execution. ManipArena comprises 20 diverse tasks across 10,812 expert trajectories emphasizing reasoning-oriented manipulation tasks requiring semantic and spatial reasoning, supports multi-level generalization through controlled out-of-distribution settings, and incorporates long-horizon mobile manipulation beyond tabletop scenarios. The framework further provides rich sensory diagnostics, including low-level motor signals, and synchronized real-to-sim environments constructed via high-quality 3D scanning. Together, these features enable fair, realistic, and reproducible evaluation for both VLA and world model approaches, providing a scalable foundation for diagnosing and advancing embodied intelligence systems. |
| title | ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation |
| topic | Robotics Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2603.28545 |