Goal Representations for Instruction Following: A Semi-Supervised Language Interface to Control

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Myers, Vivek, He, Andre, Fang, Kuan, Walke, Homer, Hansen-Estruch, Philippe, Cheng, Ching-An, Jalobeanu, Mihai, Kolobov, Andrey, Dragan, Anca, Levine, Sergey
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910776253480960
author Myers, Vivek
He, Andre
Fang, Kuan
Walke, Homer
Hansen-Estruch, Philippe
Cheng, Ching-An
Jalobeanu, Mihai
Kolobov, Andrey
Dragan, Anca
Levine, Sergey
author_facet Myers, Vivek
He, Andre
Fang, Kuan
Walke, Homer
Hansen-Estruch, Philippe
Cheng, Ching-An
Jalobeanu, Mihai
Kolobov, Andrey
Dragan, Anca
Levine, Sergey
contents Our goal is for robots to follow natural language instructions like "put the towel next to the microwave." But getting large amounts of labeled data, i.e. data that contains demonstrations of tasks labeled with the language instruction, is prohibitive. In contrast, obtaining policies that respond to image goals is much easier, because any autonomous trial or demonstration can be labeled in hindsight with its final state as the goal. In this work, we contribute a method that taps into joint image- and goal- conditioned policies with language using only a small amount of language data. Prior work has made progress on this using vision-language models or by jointly training language-goal-conditioned policies, but so far neither method has scaled effectively to real-world robot tasks without significant human annotation. Our method achieves robust performance in the real world by learning an embedding from the labeled data that aligns language not to the goal image, but rather to the desired change between the start and goal images that the instruction corresponds to. We then train a policy on this embedding: the policy benefits from all the unlabeled data, but the aligned embedding provides an interface for language to steer the policy. We show instruction following across a variety of manipulation tasks in different scenes, with generalization to language instructions outside of the labeled data. Videos and code for our approach can be found on our website: https://rail-berkeley.github.io/grif/ .
format Preprint
id arxiv_https___arxiv_org_abs_2307_00117
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Goal Representations for Instruction Following: A Semi-Supervised Language Interface to Control
Myers, Vivek
He, Andre
Fang, Kuan
Walke, Homer
Hansen-Estruch, Philippe
Cheng, Ching-An
Jalobeanu, Mihai
Kolobov, Andrey
Dragan, Anca
Levine, Sergey
Robotics
Machine Learning
Our goal is for robots to follow natural language instructions like "put the towel next to the microwave." But getting large amounts of labeled data, i.e. data that contains demonstrations of tasks labeled with the language instruction, is prohibitive. In contrast, obtaining policies that respond to image goals is much easier, because any autonomous trial or demonstration can be labeled in hindsight with its final state as the goal. In this work, we contribute a method that taps into joint image- and goal- conditioned policies with language using only a small amount of language data. Prior work has made progress on this using vision-language models or by jointly training language-goal-conditioned policies, but so far neither method has scaled effectively to real-world robot tasks without significant human annotation. Our method achieves robust performance in the real world by learning an embedding from the labeled data that aligns language not to the goal image, but rather to the desired change between the start and goal images that the instruction corresponds to. We then train a policy on this embedding: the policy benefits from all the unlabeled data, but the aligned embedding provides an interface for language to steer the policy. We show instruction following across a variety of manipulation tasks in different scenes, with generalization to language instructions outside of the labeled data. Videos and code for our approach can be found on our website: https://rail-berkeley.github.io/grif/ .
title Goal Representations for Instruction Following: A Semi-Supervised Language Interface to Control
topic Robotics
Machine Learning
url https://arxiv.org/abs/2307.00117