$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Intelligence, Physical, Black, Kevin, Brown, Noah, Darpinian, James, Dhabalia, Karan, Driess, Danny, Esmail, Adnan, Equi, Michael, Finn, Chelsea, Fusai, Niccolo, Galliker, Manuel Y., Ghosh, Dibya, Groom, Lachy, Hausman, Karol, Ichter, Brian, Jakubczak, Szymon, Jones, Tim, Ke, Liyiming, LeBlanc, Devin, Levine, Sergey, Li-Bell, Adrian, Mothukuri, Mohith, Nair, Suraj, Pertsch, Karl, Ren, Allen Z., Shi, Lucy Xiaoyang, Smith, Laura, Springenberg, Jost Tobias, Stachowicz, Kyle, Tanner, James, Vuong, Quan, Walke, Homer, Walling, Anna, Wang, Haohuan, Yu, Lili, Zhilinsky, Ury
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908332504121344
author Intelligence, Physical
Black, Kevin
Brown, Noah
Darpinian, James
Dhabalia, Karan
Driess, Danny
Esmail, Adnan
Equi, Michael
Finn, Chelsea
Fusai, Niccolo
Galliker, Manuel Y.
Ghosh, Dibya
Groom, Lachy
Hausman, Karol
Ichter, Brian
Jakubczak, Szymon
Jones, Tim
Ke, Liyiming
LeBlanc, Devin
Levine, Sergey
Li-Bell, Adrian
Mothukuri, Mohith
Nair, Suraj
Pertsch, Karl
Ren, Allen Z.
Shi, Lucy Xiaoyang
Smith, Laura
Springenberg, Jost Tobias
Stachowicz, Kyle
Tanner, James
Vuong, Quan
Walke, Homer
Walling, Anna
Wang, Haohuan
Yu, Lili
Zhilinsky, Ury
author_facet Intelligence, Physical
Black, Kevin
Brown, Noah
Darpinian, James
Dhabalia, Karan
Driess, Danny
Esmail, Adnan
Equi, Michael
Finn, Chelsea
Fusai, Niccolo
Galliker, Manuel Y.
Ghosh, Dibya
Groom, Lachy
Hausman, Karol
Ichter, Brian
Jakubczak, Szymon
Jones, Tim
Ke, Liyiming
LeBlanc, Devin
Levine, Sergey
Li-Bell, Adrian
Mothukuri, Mohith
Nair, Suraj
Pertsch, Karl
Ren, Allen Z.
Shi, Lucy Xiaoyang
Smith, Laura
Springenberg, Jost Tobias
Stachowicz, Kyle
Tanner, James
Vuong, Quan
Walke, Homer
Walling, Anna
Wang, Haohuan
Yu, Lili
Zhilinsky, Ury
contents In order for robots to be useful, they must perform practically relevant tasks in the real world, outside of the lab. While vision-language-action (VLA) models have demonstrated impressive results for end-to-end robot control, it remains an open question how far such models can generalize in the wild. We describe $π_{0.5}$, a new model based on $π_{0}$ that uses co-training on heterogeneous tasks to enable broad generalization. $π_{0.5}$\ uses data from multiple robots, high-level semantic prediction, web data, and other sources to enable broadly generalizable real-world robotic manipulation. Our system uses a combination of co-training and hybrid multi-modal examples that combine image observations, language commands, object detections, semantic subtask prediction, and low-level actions. Our experiments show that this kind of knowledge transfer is essential for effective generalization, and we demonstrate for the first time that an end-to-end learning-enabled robotic system can perform long-horizon and dexterous manipulation skills, such as cleaning a kitchen or bedroom, in entirely new homes.
format Preprint
id arxiv_https___arxiv_org_abs_2504_16054
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle $π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
Intelligence, Physical
Black, Kevin
Brown, Noah
Darpinian, James
Dhabalia, Karan
Driess, Danny
Esmail, Adnan
Equi, Michael
Finn, Chelsea
Fusai, Niccolo
Galliker, Manuel Y.
Ghosh, Dibya
Groom, Lachy
Hausman, Karol
Ichter, Brian
Jakubczak, Szymon
Jones, Tim
Ke, Liyiming
LeBlanc, Devin
Levine, Sergey
Li-Bell, Adrian
Mothukuri, Mohith
Nair, Suraj
Pertsch, Karl
Ren, Allen Z.
Shi, Lucy Xiaoyang
Smith, Laura
Springenberg, Jost Tobias
Stachowicz, Kyle
Tanner, James
Vuong, Quan
Walke, Homer
Walling, Anna
Wang, Haohuan
Yu, Lili
Zhilinsky, Ury
Machine Learning
Robotics
In order for robots to be useful, they must perform practically relevant tasks in the real world, outside of the lab. While vision-language-action (VLA) models have demonstrated impressive results for end-to-end robot control, it remains an open question how far such models can generalize in the wild. We describe $π_{0.5}$, a new model based on $π_{0}$ that uses co-training on heterogeneous tasks to enable broad generalization. $π_{0.5}$\ uses data from multiple robots, high-level semantic prediction, web data, and other sources to enable broadly generalizable real-world robotic manipulation. Our system uses a combination of co-training and hybrid multi-modal examples that combine image observations, language commands, object detections, semantic subtask prediction, and low-level actions. Our experiments show that this kind of knowledge transfer is essential for effective generalization, and we demonstrate for the first time that an end-to-end learning-enabled robotic system can perform long-horizon and dexterous manipulation skills, such as cleaning a kitchen or bedroom, in entirely new homes.
title $π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
topic Machine Learning
Robotics
url https://arxiv.org/abs/2504.16054