3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ze, Yanjie, Zhang, Gu, Zhang, Kangning, Hu, Chenyuan, Wang, Muhan, Xu, Huazhe
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912047677046784
author Ze, Yanjie
Zhang, Gu
Zhang, Kangning
Hu, Chenyuan
Wang, Muhan
Xu, Huazhe
author_facet Ze, Yanjie
Zhang, Gu
Zhang, Kangning
Hu, Chenyuan
Wang, Muhan
Xu, Huazhe
contents Imitation learning provides an efficient way to teach robots dexterous skills; however, learning complex skills robustly and generalizablely usually consumes large amounts of human demonstrations. To tackle this challenging problem, we present 3D Diffusion Policy (DP3), a novel visual imitation learning approach that incorporates the power of 3D visual representations into diffusion policies, a class of conditional action generative models. The core design of DP3 is the utilization of a compact 3D visual representation, extracted from sparse point clouds with an efficient point encoder. In our experiments involving 72 simulation tasks, DP3 successfully handles most tasks with just 10 demonstrations and surpasses baselines with a 24.2% relative improvement. In 4 real robot tasks, DP3 demonstrates precise control with a high success rate of 85%, given only 40 demonstrations of each task, and shows excellent generalization abilities in diverse aspects, including space, viewpoint, appearance, and instance. Interestingly, in real robot experiments, DP3 rarely violates safety requirements, in contrast to baseline methods which frequently do, necessitating human intervention. Our extensive evaluation highlights the critical importance of 3D representations in real-world robot learning. Videos, code, and data are available on https://3d-diffusion-policy.github.io .
format Preprint
id arxiv_https___arxiv_org_abs_2403_03954
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations
Ze, Yanjie
Zhang, Gu
Zhang, Kangning
Hu, Chenyuan
Wang, Muhan
Xu, Huazhe
Robotics
Computer Vision and Pattern Recognition
Machine Learning
Imitation learning provides an efficient way to teach robots dexterous skills; however, learning complex skills robustly and generalizablely usually consumes large amounts of human demonstrations. To tackle this challenging problem, we present 3D Diffusion Policy (DP3), a novel visual imitation learning approach that incorporates the power of 3D visual representations into diffusion policies, a class of conditional action generative models. The core design of DP3 is the utilization of a compact 3D visual representation, extracted from sparse point clouds with an efficient point encoder. In our experiments involving 72 simulation tasks, DP3 successfully handles most tasks with just 10 demonstrations and surpasses baselines with a 24.2% relative improvement. In 4 real robot tasks, DP3 demonstrates precise control with a high success rate of 85%, given only 40 demonstrations of each task, and shows excellent generalization abilities in diverse aspects, including space, viewpoint, appearance, and instance. Interestingly, in real robot experiments, DP3 rarely violates safety requirements, in contrast to baseline methods which frequently do, necessitating human intervention. Our extensive evaluation highlights the critical importance of 3D representations in real-world robot learning. Videos, code, and data are available on https://3d-diffusion-policy.github.io .
title 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations
topic Robotics
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2403.03954