H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bi, Hongzhe, Wu, Lingxuan, Lin, Tianwei, Tan, Hengkai, Su, Zhizhong, Su, Hang, Zhu, Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913969296375808
author Bi, Hongzhe
Wu, Lingxuan
Lin, Tianwei
Tan, Hengkai
Su, Zhizhong
Su, Hang
Zhu, Jun
author_facet Bi, Hongzhe
Wu, Lingxuan
Lin, Tianwei
Tan, Hengkai
Su, Zhizhong
Su, Hang
Zhu, Jun
contents Imitation learning for robotic manipulation faces a fundamental challenge: the scarcity of large-scale, high-quality robot demonstration data. Recent robotic foundation models often pre-train on cross-embodiment robot datasets to increase data scale, while they face significant limitations as the diverse morphologies and action spaces across different robot embodiments make unified training challenging. In this paper, we present H-RDT (Human to Robotics Diffusion Transformer), a novel approach that leverages human manipulation data to enhance robot manipulation capabilities. Our key insight is that large-scale egocentric human manipulation videos with paired 3D hand pose annotations provide rich behavioral priors that capture natural manipulation strategies and can benefit robotic policy learning. We introduce a two-stage training paradigm: (1) pre-training on large-scale egocentric human manipulation data, and (2) cross-embodiment fine-tuning on robot-specific data with modular action encoders and decoders. Built on a diffusion transformer architecture with 2B parameters, H-RDT uses flow matching to model complex action distributions. Extensive evaluations encompassing both simulation and real-world experiments, single-task and multitask scenarios, as well as few-shot learning and robustness assessments, demonstrate that H-RDT outperforms training from scratch and existing state-of-the-art methods, including Pi0 and RDT, achieving significant improvements of 13.9% and 40.5% over training from scratch in simulation and real-world experiments, respectively. The results validate our core hypothesis that human manipulation data can serve as a powerful foundation for learning bimanual robotic manipulation policies.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23523
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
Bi, Hongzhe
Wu, Lingxuan
Lin, Tianwei
Tan, Hengkai
Su, Zhizhong
Su, Hang
Zhu, Jun
Robotics
Computer Vision and Pattern Recognition
Machine Learning
Imitation learning for robotic manipulation faces a fundamental challenge: the scarcity of large-scale, high-quality robot demonstration data. Recent robotic foundation models often pre-train on cross-embodiment robot datasets to increase data scale, while they face significant limitations as the diverse morphologies and action spaces across different robot embodiments make unified training challenging. In this paper, we present H-RDT (Human to Robotics Diffusion Transformer), a novel approach that leverages human manipulation data to enhance robot manipulation capabilities. Our key insight is that large-scale egocentric human manipulation videos with paired 3D hand pose annotations provide rich behavioral priors that capture natural manipulation strategies and can benefit robotic policy learning. We introduce a two-stage training paradigm: (1) pre-training on large-scale egocentric human manipulation data, and (2) cross-embodiment fine-tuning on robot-specific data with modular action encoders and decoders. Built on a diffusion transformer architecture with 2B parameters, H-RDT uses flow matching to model complex action distributions. Extensive evaluations encompassing both simulation and real-world experiments, single-task and multitask scenarios, as well as few-shot learning and robustness assessments, demonstrate that H-RDT outperforms training from scratch and existing state-of-the-art methods, including Pi0 and RDT, achieving significant improvements of 13.9% and 40.5% over training from scratch in simulation and real-world experiments, respectively. The results validate our core hypothesis that human manipulation data can serve as a powerful foundation for learning bimanual robotic manipulation policies.
title H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
topic Robotics
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2507.23523