ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sun, Yu, Cao, Meng, Yang, Ping, Xu, Rongtao, Yan, Yunxiao, Xu, Runze, Ma, Liang, Gan, Roy, Zhai, Andy, Chen, Qingxuan, Xu, Zunnan, Wang, Hao, Yu, Jincheng, Liang, Lucy, Wang, Qian, Laptev, Ivan, Reid, Ian D, Liang, Xiaodan
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911553720156160
author Sun, Yu
Cao, Meng
Yang, Ping
Xu, Rongtao
Yan, Yunxiao
Xu, Runze
Ma, Liang
Gan, Roy
Zhai, Andy
Chen, Qingxuan
Xu, Zunnan
Wang, Hao
Yu, Jincheng
Liang, Lucy
Wang, Qian
Laptev, Ivan
Reid, Ian D
Liang, Xiaodan
author_facet Sun, Yu
Cao, Meng
Yang, Ping
Xu, Rongtao
Yan, Yunxiao
Xu, Runze
Ma, Liang
Gan, Roy
Zhai, Andy
Chen, Qingxuan
Xu, Zunnan
Wang, Hao
Yu, Jincheng
Liang, Lucy
Wang, Qian
Laptev, Ivan
Reid, Ian D
Liang, Xiaodan
contents Vision-Language-Action (VLA) models and world models have recently emerged as promising paradigms for general-purpose robotic intelligence, yet their progress is hindered by the lack of reliable evaluation protocols that reflect real-world deployment. Existing benchmarks are largely simulator-centric, which provide controllability but fail to capture the reality gap caused by perception noise, complex contact dynamics, hardware constraints, and system latency. Moreover, fragmented real-world evaluations across different robot platforms prevent fair and reproducible comparison. To address these challenges, we introduce ManipArena, a standardized evaluation framework designed to bridge simulation and real-world execution. ManipArena comprises 20 diverse tasks across 10,812 expert trajectories emphasizing reasoning-oriented manipulation tasks requiring semantic and spatial reasoning, supports multi-level generalization through controlled out-of-distribution settings, and incorporates long-horizon mobile manipulation beyond tabletop scenarios. The framework further provides rich sensory diagnostics, including low-level motor signals, and synchronized real-to-sim environments constructed via high-quality 3D scanning. Together, these features enable fair, realistic, and reproducible evaluation for both VLA and world model approaches, providing a scalable foundation for diagnosing and advancing embodied intelligence systems.
format Preprint
id arxiv_https___arxiv_org_abs_2603_28545
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation
Sun, Yu
Cao, Meng
Yang, Ping
Xu, Rongtao
Yan, Yunxiao
Xu, Runze
Ma, Liang
Gan, Roy
Zhai, Andy
Chen, Qingxuan
Xu, Zunnan
Wang, Hao
Yu, Jincheng
Liang, Lucy
Wang, Qian
Laptev, Ivan
Reid, Ian D
Liang, Xiaodan
Robotics
Computer Vision and Pattern Recognition
Vision-Language-Action (VLA) models and world models have recently emerged as promising paradigms for general-purpose robotic intelligence, yet their progress is hindered by the lack of reliable evaluation protocols that reflect real-world deployment. Existing benchmarks are largely simulator-centric, which provide controllability but fail to capture the reality gap caused by perception noise, complex contact dynamics, hardware constraints, and system latency. Moreover, fragmented real-world evaluations across different robot platforms prevent fair and reproducible comparison. To address these challenges, we introduce ManipArena, a standardized evaluation framework designed to bridge simulation and real-world execution. ManipArena comprises 20 diverse tasks across 10,812 expert trajectories emphasizing reasoning-oriented manipulation tasks requiring semantic and spatial reasoning, supports multi-level generalization through controlled out-of-distribution settings, and incorporates long-horizon mobile manipulation beyond tabletop scenarios. The framework further provides rich sensory diagnostics, including low-level motor signals, and synchronized real-to-sim environments constructed via high-quality 3D scanning. Together, these features enable fair, realistic, and reproducible evaluation for both VLA and world model approaches, providing a scalable foundation for diagnosing and advancing embodied intelligence systems.
title ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.28545