The Ingredients for Robotic Diffusion Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dasari, Sudeep, Mees, Oier, Zhao, Sebastian, Srirama, Mohan Kumar, Levine, Sergey
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916437587656704
author Dasari, Sudeep
Mees, Oier
Zhao, Sebastian
Srirama, Mohan Kumar
Levine, Sergey
author_facet Dasari, Sudeep
Mees, Oier
Zhao, Sebastian
Srirama, Mohan Kumar
Levine, Sergey
contents In recent years roboticists have achieved remarkable progress in solving increasingly general tasks on dexterous robotic hardware by leveraging high capacity Transformer network architectures and generative diffusion models. Unfortunately, combining these two orthogonal improvements has proven surprisingly difficult, since there is no clear and well-understood process for making important design choices. In this paper, we identify, study and improve key architectural design decisions for high-capacity diffusion transformer policies. The resulting models can efficiently solve diverse tasks on multiple robot embodiments, without the excruciating pain of per-setup hyper-parameter tuning. By combining the results of our investigation with our improved model components, we are able to present a novel architecture, named \method, that significantly outperforms the state of the art in solving long-horizon ($1500+$ time-steps) dexterous tasks on a bi-manual ALOHA robot. In addition, we find that our policies show improved scaling performance when trained on 10 hours of highly multi-modal, language annotated ALOHA demonstration data. We hope this work will open the door for future robot learning techniques that leverage the efficiency of generative diffusion modeling with the scalability of large scale transformer architectures. Code, robot dataset, and videos are available at: https://dit-policy.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2410_10088
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Ingredients for Robotic Diffusion Transformers
Dasari, Sudeep
Mees, Oier
Zhao, Sebastian
Srirama, Mohan Kumar
Levine, Sergey
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
In recent years roboticists have achieved remarkable progress in solving increasingly general tasks on dexterous robotic hardware by leveraging high capacity Transformer network architectures and generative diffusion models. Unfortunately, combining these two orthogonal improvements has proven surprisingly difficult, since there is no clear and well-understood process for making important design choices. In this paper, we identify, study and improve key architectural design decisions for high-capacity diffusion transformer policies. The resulting models can efficiently solve diverse tasks on multiple robot embodiments, without the excruciating pain of per-setup hyper-parameter tuning. By combining the results of our investigation with our improved model components, we are able to present a novel architecture, named \method, that significantly outperforms the state of the art in solving long-horizon ($1500+$ time-steps) dexterous tasks on a bi-manual ALOHA robot. In addition, we find that our policies show improved scaling performance when trained on 10 hours of highly multi-modal, language annotated ALOHA demonstration data. We hope this work will open the door for future robot learning techniques that leverage the efficiency of generative diffusion modeling with the scalability of large scale transformer architectures. Code, robot dataset, and videos are available at: https://dit-policy.github.io
title The Ingredients for Robotic Diffusion Transformers
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2410.10088