Latent Action Pretraining from Videos

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ye, Seonghyeon, Jang, Joel, Jeon, Byeongguk, Joo, Sejune, Yang, Jianwei, Peng, Baolin, Mandlekar, Ajay, Tan, Reuben, Chao, Yu-Wei, Lin, Bill Yuchen, Liden, Lars, Lee, Kimin, Gao, Jianfeng, Zettlemoyer, Luke, Fox, Dieter, Seo, Minjoon
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910944333922304
author Ye, Seonghyeon
Jang, Joel
Jeon, Byeongguk
Joo, Sejune
Yang, Jianwei
Peng, Baolin
Mandlekar, Ajay
Tan, Reuben
Chao, Yu-Wei
Lin, Bill Yuchen
Liden, Lars
Lee, Kimin
Gao, Jianfeng
Zettlemoyer, Luke
Fox, Dieter
Seo, Minjoon
author_facet Ye, Seonghyeon
Jang, Joel
Jeon, Byeongguk
Joo, Sejune
Yang, Jianwei
Peng, Baolin
Mandlekar, Ajay
Tan, Reuben
Chao, Yu-Wei
Lin, Bill Yuchen
Liden, Lars
Lee, Kimin
Gao, Jianfeng
Zettlemoyer, Luke
Fox, Dieter
Seo, Minjoon
contents We introduce Latent Action Pretraining for general Action models (LAPA), an unsupervised method for pretraining Vision-Language-Action (VLA) models without ground-truth robot action labels. Existing Vision-Language-Action models require action labels typically collected by human teleoperators during pretraining, which significantly limits possible data sources and scale. In this work, we propose a method to learn from internet-scale videos that do not have robot action labels. We first train an action quantization model leveraging VQ-VAE-based objective to learn discrete latent actions between image frames, then pretrain a latent VLA model to predict these latent actions from observations and task descriptions, and finally finetune the VLA on small-scale robot manipulation data to map from latent to robot actions. Experimental results demonstrate that our method significantly outperforms existing techniques that train robot manipulation policies from large-scale videos. Furthermore, it outperforms the state-of-the-art VLA model trained with robotic action labels on real-world manipulation tasks that require language conditioning, generalization to unseen objects, and semantic generalization to unseen instructions. Training only on human manipulation videos also shows positive transfer, opening up the potential for leveraging web-scale data for robotics foundation model.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11758
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Latent Action Pretraining from Videos
Ye, Seonghyeon
Jang, Joel
Jeon, Byeongguk
Joo, Sejune
Yang, Jianwei
Peng, Baolin
Mandlekar, Ajay
Tan, Reuben
Chao, Yu-Wei
Lin, Bill Yuchen
Liden, Lars
Lee, Kimin
Gao, Jianfeng
Zettlemoyer, Luke
Fox, Dieter
Seo, Minjoon
Robotics
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
We introduce Latent Action Pretraining for general Action models (LAPA), an unsupervised method for pretraining Vision-Language-Action (VLA) models without ground-truth robot action labels. Existing Vision-Language-Action models require action labels typically collected by human teleoperators during pretraining, which significantly limits possible data sources and scale. In this work, we propose a method to learn from internet-scale videos that do not have robot action labels. We first train an action quantization model leveraging VQ-VAE-based objective to learn discrete latent actions between image frames, then pretrain a latent VLA model to predict these latent actions from observations and task descriptions, and finally finetune the VLA on small-scale robot manipulation data to map from latent to robot actions. Experimental results demonstrate that our method significantly outperforms existing techniques that train robot manipulation policies from large-scale videos. Furthermore, it outperforms the state-of-the-art VLA model trained with robotic action labels on real-world manipulation tasks that require language conditioning, generalization to unseen objects, and semantic generalization to unseen instructions. Training only on human manipulation videos also shows positive transfer, opening up the potential for leveraging web-scale data for robotics foundation model.
title Latent Action Pretraining from Videos
topic Robotics
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2410.11758