Trustworthy Evaluation of Robotic Manipulation: A New Benchmark and AutoEval Methods
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Mengyuan, Sheng, Juyi, Li, Peiming, Wang, Ziyi, Xu, Tianming, Xu, Tiantian, Liu, Hong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MP1: MeanFlow Tames Policy Learning in 1-step for Robotic Manipulation
by: Sheng, Juyi, et al.
Published: (2025)
by: Sheng, Juyi, et al.
Published: (2025)
AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
by: Zhou, Zhiyuan, et al.
Published: (2025)
by: Zhou, Zhiyuan, et al.
Published: (2025)
GPA-RAM: Grasp-Pretraining Augmented Robotic Attention Mamba for Spatial Task Learning
by: Sheng, Juyi, et al.
Published: (2025)
by: Sheng, Juyi, et al.
Published: (2025)
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
Lens Privacy Sealing: A New Benchmark and Method for Physical Privacy-Preserving Action Recognition
by: Liu, Mengyuan, et al.
Published: (2026)
by: Liu, Mengyuan, et al.
Published: (2026)
AutoEval: A Practical Framework for Autonomous Evaluation of Mobile Agents
by: Sun, Jiahui, et al.
Published: (2025)
by: Sun, Jiahui, et al.
Published: (2025)
RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation
by: Jiang, Feng, et al.
Published: (2026)
by: Jiang, Feng, et al.
Published: (2026)
AutoEval Done Right: Using Synthetic Data for Model Evaluation
by: Boyeau, Pierre, et al.
Published: (2024)
by: Boyeau, Pierre, et al.
Published: (2024)
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
by: Li, Huiqiong, et al.
Published: (2026)
by: Li, Huiqiong, et al.
Published: (2026)
FMB: a Functional Manipulation Benchmark for Generalizable Robotic Learning
by: Luo, Jianlan, et al.
Published: (2024)
by: Luo, Jianlan, et al.
Published: (2024)
RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation
by: Wang, Yi Ru, et al.
Published: (2025)
by: Wang, Yi Ru, et al.
Published: (2025)
Adaptive Prediction-Powered AutoEval with Reliability and Efficiency Guarantees
by: Park, Sangwoo, et al.
Published: (2025)
by: Park, Sangwoo, et al.
Published: (2025)
Toward Visually Realistic Simulation: A Benchmark for Evaluating Robot Manipulation in Simulation
by: Zhu, Yixin, et al.
Published: (2026)
by: Zhu, Yixin, et al.
Published: (2026)
FORGE-Tree: Diffusion-Forcing Tree Search for Long-Horizon Robot Manipulation
by: Huang, Yanjia, et al.
Published: (2025)
by: Huang, Yanjia, et al.
Published: (2025)
HERMES: Human-to-Robot Embodied Learning from Multi-Source Motion Data for Mobile Dexterous Manipulation
by: Yuan, Zhecheng, et al.
Published: (2025)
by: Yuan, Zhecheng, et al.
Published: (2025)
Diffusion Models for Robotic Manipulation: A Survey
by: Wolf, Rosa, et al.
Published: (2025)
by: Wolf, Rosa, et al.
Published: (2025)
Goal State Generation for Robotic Manipulation Based on Linguistically Guided Hybrid Gaussian Diffusion
by: Xu, Yichen, et al.
Published: (2024)
by: Xu, Yichen, et al.
Published: (2024)
THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
by: Pumacay, Wilbert, et al.
Published: (2024)
by: Pumacay, Wilbert, et al.
Published: (2024)
AutoBio: A Simulation and Benchmark for Robotic Automation in Digital Biology Laboratory
by: Lan, Zhiqian, et al.
Published: (2025)
by: Lan, Zhiqian, et al.
Published: (2025)
RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic Manipulation
by: Lin, Sixu, et al.
Published: (2026)
by: Lin, Sixu, et al.
Published: (2026)
WorldEval: World Model as Real-World Robot Policies Evaluator
by: Li, Yaxuan, et al.
Published: (2025)
by: Li, Yaxuan, et al.
Published: (2025)
AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
by: Xing, Shuo, et al.
Published: (2024)
by: Xing, Shuo, et al.
Published: (2024)
Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation
by: Qi, Yu, et al.
Published: (2025)
by: Qi, Yu, et al.
Published: (2025)
Visual Robotic Manipulation with Depth-Aware Pretraining
by: Wang, Wanying, et al.
Published: (2024)
by: Wang, Wanying, et al.
Published: (2024)
SafeManip: A Property-Driven Benchmark for Temporal Safety Evaluation in Robotic Manipulation
by: Huang, Chengyue, et al.
Published: (2026)
by: Huang, Chengyue, et al.
Published: (2026)
AffordSim: A Scalable Data Generator and Benchmark for Affordance-Aware Robotic Manipulation
by: Li, Mingyang, et al.
Published: (2026)
by: Li, Mingyang, et al.
Published: (2026)
RoboBERT: An End-to-end Multimodal Robotic Manipulation Model
by: Wang, Sicheng, et al.
Published: (2025)
by: Wang, Sicheng, et al.
Published: (2025)
RMBench: Memory-Dependent Robotic Manipulation Benchmark with Insights into Policy Design
by: Chen, Tianxing, et al.
Published: (2026)
by: Chen, Tianxing, et al.
Published: (2026)
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
by: Han, Songhao, et al.
Published: (2025)
by: Han, Songhao, et al.
Published: (2025)
The Developments and Challenges towards Dexterous and Embodied Robotic Manipulation: A Survey
by: Li, Gaofeng, et al.
Published: (2025)
by: Li, Gaofeng, et al.
Published: (2025)
RoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation
by: Wu, Kun, et al.
Published: (2024)
by: Wu, Kun, et al.
Published: (2024)
ManiPose: A Comprehensive Benchmark for Pose-aware Object Manipulation in Robotics
by: Yu, Qiaojun, et al.
Published: (2024)
by: Yu, Qiaojun, et al.
Published: (2024)
Query-Centric Diffusion Policy for Generalizable Robotic Assembly
by: Xu, Ziyi, et al.
Published: (2025)
by: Xu, Ziyi, et al.
Published: (2025)
FLAME: A Federated Learning Benchmark for Robotic Manipulation
by: Betran, Santiago Bou, et al.
Published: (2025)
by: Betran, Santiago Bou, et al.
Published: (2025)
DexImit: Learning Bimanual Dexterous Manipulation from Monocular Human Videos
by: Mu, Juncheng, et al.
Published: (2026)
by: Mu, Juncheng, et al.
Published: (2026)
GazeVLA: Learning Human Intention for Robotic Manipulation
by: Li, Chengyang, et al.
Published: (2026)
by: Li, Chengyang, et al.
Published: (2026)
3D Affordance Keypoint Detection for Robotic Manipulation
by: Liu, Zhiyang, et al.
Published: (2025)
by: Liu, Zhiyang, et al.
Published: (2025)
From Reaction to Anticipation: Proactive Failure Recovery through Agentic Task Graph for Robotic Manipulation
by: Xu, Sheng, et al.
Published: (2026)
by: Xu, Sheng, et al.
Published: (2026)
ClickDiff: Click to Induce Semantic Contact Map for Controllable Grasp Generation with Diffusion Models
by: Li, Peiming, et al.
Published: (2024)
by: Li, Peiming, et al.
Published: (2024)
ActivePose: Active 6D Object Pose Estimation and Tracking for Robotic Manipulation
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
Similar Items
-
MP1: MeanFlow Tames Policy Learning in 1-step for Robotic Manipulation
by: Sheng, Juyi, et al.
Published: (2025) -
AutoEval: Autonomous Evaluation of Generalist Robot Manipulation Policies in the Real World
by: Zhou, Zhiyuan, et al.
Published: (2025) -
GPA-RAM: Grasp-Pretraining Augmented Robotic Attention Mamba for Spatial Task Learning
by: Sheng, Juyi, et al.
Published: (2025) -
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
by: Wang, Ziyi, et al.
Published: (2025) -
Lens Privacy Sealing: A New Benchmark and Method for Physical Privacy-Preserving Action Recognition
by: Liu, Mengyuan, et al.
Published: (2026)