ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Yu, Cao, Meng, Yang, Ping, Xu, Rongtao, Yan, Yunxiao, Xu, Runze, Ma, Liang, Gan, Roy, Zhai, Andy, Chen, Qingxuan, Xu, Zunnan, Wang, Hao, Yu, Jincheng, Liang, Lucy, Wang, Qian, Laptev, Ivan, Reid, Ian D, Liang, Xiaodan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RoBridge: A Hierarchical Architecture Bridging Cognition and Execution for General Robotic Manipulation
by: Zhang, Kaidong, et al.
Published: (2025)
by: Zhang, Kaidong, et al.
Published: (2025)
GLaD: Geometric Latent Distillation for Vision-Language-Action Models
by: Guo, Minghao, et al.
Published: (2025)
by: Guo, Minghao, et al.
Published: (2025)
InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models
by: Yan, Yu, et al.
Published: (2024)
by: Yan, Yu, et al.
Published: (2024)
3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
by: Xu, Rongtao, et al.
Published: (2025)
by: Xu, Rongtao, et al.
Published: (2025)
ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation
by: Xu, Zitong, et al.
Published: (2025)
by: Xu, Zitong, et al.
Published: (2025)
HeteroGenManip: Generalizable Manipulation For Heterogeneous Object Interactions
by: Shen, Zhenhao, et al.
Published: (2026)
by: Shen, Zhenhao, et al.
Published: (2026)
Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
by: Cao, Meng, et al.
Published: (2025)
by: Cao, Meng, et al.
Published: (2025)
UniManip: General-Purpose Zero-Shot Robotic Manipulation with Agentic Operational Graph
by: Liu, Haichao, et al.
Published: (2026)
by: Liu, Haichao, et al.
Published: (2026)
CycleManip: Enabling Cyclic Task Manipulation via Effective Historical Perception and Understanding
by: Wei, Yi-Lin, et al.
Published: (2025)
by: Wei, Yi-Lin, et al.
Published: (2025)
MoManipVLA: Transferring Vision-language-action Models for General Mobile Manipulation
by: Wu, Zhenyu, et al.
Published: (2025)
by: Wu, Zhenyu, et al.
Published: (2025)
Complementary Information Guided Occupancy Prediction via Multi-Level Representation Fusion
by: Xu, Rongtao, et al.
Published: (2025)
by: Xu, Rongtao, et al.
Published: (2025)
PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly
by: Ma, Liang, et al.
Published: (2025)
by: Ma, Liang, et al.
Published: (2025)
Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
by: Chen, Kehan, et al.
Published: (2024)
by: Chen, Kehan, et al.
Published: (2024)
AdaManip: Adaptive Articulated Object Manipulation Environments and Policy Learning
by: Wang, Yuanfei, et al.
Published: (2025)
by: Wang, Yuanfei, et al.
Published: (2025)
iManip: Skill-Incremental Learning for Robotic Manipulation
by: Zheng, Zexin, et al.
Published: (2025)
by: Zheng, Zexin, et al.
Published: (2025)
Structured Preference Optimization for Vision-Language Long-Horizon Task Planning
by: Liang, Xiwen, et al.
Published: (2025)
by: Liang, Xiwen, et al.
Published: (2025)
World2Act: Latent Action Post-Training from World Model Dynamics
by: Vuong, An Dinh, et al.
Published: (2026)
by: Vuong, An Dinh, et al.
Published: (2026)
ArtiSG: Functional 3D Scene Graph Construction via Human-demonstrated Articulated Objects Manipulation
by: Gu, Qiuyi, et al.
Published: (2025)
by: Gu, Qiuyi, et al.
Published: (2025)
ActionSink: Toward Precise Robot Manipulation with Dynamic Integration of Action Flow
by: Guo, Shanshan, et al.
Published: (2025)
by: Guo, Shanshan, et al.
Published: (2025)
A0: An Affordance-Aware Hierarchical Model for General Robotic Manipulation
by: Xu, Rongtao, et al.
Published: (2025)
by: Xu, Rongtao, et al.
Published: (2025)
MentalManip: A Dataset For Fine-grained Analysis of Mental Manipulation in Conversations
by: Wang, Yuxin, et al.
Published: (2024)
by: Wang, Yuxin, et al.
Published: (2024)
ad-trait: A Fast and Flexible Automatic Differentiation Library in Rust
by: Liang, Chen, et al.
Published: (2025)
by: Liang, Chen, et al.
Published: (2025)
Figure 1 from: Xu S-Z, Xu H, Gan Q-L, Li Z-Y (2025) Lysimachia speciosa (Primulaceae), a new species from Central China. PhytoKeys 263: 209-214. https://doi.org/10.3897/phytokeys.263.139659
by: Xu, Song-Zhi, et al.
Published: (2025)
by: Xu, Song-Zhi, et al.
Published: (2025)
MaPPER: Multimodal Prior-guided Parameter Efficient Tuning for Referring Expression Comprehension
by: Liu, Ting, et al.
Published: (2024)
by: Liu, Ting, et al.
Published: (2024)
RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation
by: Han, Mingfei, et al.
Published: (2024)
by: Han, Mingfei, et al.
Published: (2024)
EchoTrail-GUI: Building Actionable Memory for GUI Agents via Critic-Guided Self-Exploration
by: Li, Runze, et al.
Published: (2025)
by: Li, Runze, et al.
Published: (2025)
ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models
by: Song, Zirui, et al.
Published: (2025)
by: Song, Zirui, et al.
Published: (2025)
Reading the Cell, Designing the Cure: Perturbation-Conditioned Molecular Diffusion for Function-Oriented Drug Design
by: Xu, Ziyu, et al.
Published: (2026)
by: Xu, Ziyu, et al.
Published: (2026)
DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance
by: Wang, Cong, et al.
Published: (2023)
by: Wang, Cong, et al.
Published: (2023)
AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
by: Zhou, Zhiyuan, et al.
Published: (2025)
by: Zhou, Zhiyuan, et al.
Published: (2025)
Toward Generalist Neural Motion Planners for Robotic Manipulators: Challenges and Opportunities
by: Soleymanzadeh, Davood, et al.
Published: (2026)
by: Soleymanzadeh, Davood, et al.
Published: (2026)
FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Model
by: Zhou, Jun, et al.
Published: (2025)
by: Zhou, Jun, et al.
Published: (2025)
CAE-DFKD: Bridging the Transferability Gap in Data-Free Knowledge Distillation
by: Zhang, Zherui, et al.
Published: (2025)
by: Zhang, Zherui, et al.
Published: (2025)
PIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic Manipulation
by: Zhang, Kaidong, et al.
Published: (2024)
by: Zhang, Kaidong, et al.
Published: (2024)
ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation
by: Zhao, Enyu, et al.
Published: (2025)
by: Zhao, Enyu, et al.
Published: (2025)
Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
Artificial Intelligence-Assistant Cardiotocography: Unified Model for Signal Reconstruction, Fetal Heart Rate Analysis, and Variability Assessment
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCo
by: Wang, Hanwen, et al.
Published: (2026)
by: Wang, Hanwen, et al.
Published: (2026)
Explaining Black-box Language Models with Knowledge Probing Systems: A Post-hoc Explanation Perspective
by: Zhao, Yunxiao, et al.
Published: (2025)
by: Zhao, Yunxiao, et al.
Published: (2025)
Similar Items
-
RoBridge: A Hierarchical Architecture Bridging Cognition and Execution for General Robotic Manipulation
by: Zhang, Kaidong, et al.
Published: (2025) -
GLaD: Geometric Latent Distillation for Vision-Language-Action Models
by: Guo, Minghao, et al.
Published: (2025) -
InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models
by: Yan, Yu, et al.
Published: (2024) -
3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
by: Xu, Rongtao, et al.
Published: (2025) -
ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation
by: Xu, Zitong, et al.
Published: (2025)