This&That: Language-Gesture Controlled Video Generation for Robot Planning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Boyang, Sridhar, Nikhil, Feng, Chao, Van der Merwe, Mark, Fishman, Adam, Fazeli, Nima, Park, Jeong Joon
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918022591021056
author Wang, Boyang
Sridhar, Nikhil
Feng, Chao
Van der Merwe, Mark
Fishman, Adam
Fazeli, Nima
Park, Jeong Joon
author_facet Wang, Boyang
Sridhar, Nikhil
Feng, Chao
Van der Merwe, Mark
Fishman, Adam
Fazeli, Nima
Park, Jeong Joon
contents Clear, interpretable instructions are invaluable when attempting any complex task. Good instructions help to clarify the task and even anticipate the steps needed to solve it. In this work, we propose a robot learning framework for communicating, planning, and executing a wide range of tasks, dubbed This&That. This&That solves general tasks by leveraging video generative models, which, through training on internet-scale data, contain rich physical and semantic context. In this work, we tackle three fundamental challenges in video-based planning: 1) unambiguous task communication with simple human instructions, 2) controllable video generation that respects user intent, and 3) translating visual plans into robot actions. This&That uses language-gesture conditioning to generate video predictions, as a succinct and unambiguous alternative to existing language-only methods, especially in complex and uncertain environments. These video predictions are then fed into a behavior cloning architecture dubbed Diffusion Video to Action (DiVA), which outperforms prior state-of-the-art behavior cloning and video-based planning methods by substantial margins.
format Preprint
id arxiv_https___arxiv_org_abs_2407_05530
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle This&That: Language-Gesture Controlled Video Generation for Robot Planning
Wang, Boyang
Sridhar, Nikhil
Feng, Chao
Van der Merwe, Mark
Fishman, Adam
Fazeli, Nima
Park, Jeong Joon
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Clear, interpretable instructions are invaluable when attempting any complex task. Good instructions help to clarify the task and even anticipate the steps needed to solve it. In this work, we propose a robot learning framework for communicating, planning, and executing a wide range of tasks, dubbed This&That. This&That solves general tasks by leveraging video generative models, which, through training on internet-scale data, contain rich physical and semantic context. In this work, we tackle three fundamental challenges in video-based planning: 1) unambiguous task communication with simple human instructions, 2) controllable video generation that respects user intent, and 3) translating visual plans into robot actions. This&That uses language-gesture conditioning to generate video predictions, as a succinct and unambiguous alternative to existing language-only methods, especially in complex and uncertain environments. These video predictions are then fed into a behavior cloning architecture dubbed Diffusion Video to Action (DiVA), which outperforms prior state-of-the-art behavior cloning and video-based planning methods by substantial margins.
title This&That: Language-Gesture Controlled Video Generation for Robot Planning
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.05530