One-Shot Imitation under Mismatched Execution

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kedia, Kushal, Dan, Prithwish, Chao, Angela, Pace, Maximus Adrian, Choudhury, Sanjiban
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913764852367360
author Kedia, Kushal
Dan, Prithwish
Chao, Angela
Pace, Maximus Adrian
Choudhury, Sanjiban
author_facet Kedia, Kushal
Dan, Prithwish
Chao, Angela
Pace, Maximus Adrian
Choudhury, Sanjiban
contents Human demonstrations as prompts are a powerful way to program robots to do long-horizon manipulation tasks. However, translating these demonstrations into robot-executable actions presents significant challenges due to execution mismatches in movement styles and physical capabilities. Existing methods for human-robot translation either depend on paired data, which is infeasible to scale, or rely heavily on frame-level visual similarities that often break down in practice. To address these challenges, we propose RHyME, a novel framework that automatically pairs human and robot trajectories using sequence-level optimal transport cost functions. Given long-horizon robot demonstrations, RHyME synthesizes semantically equivalent human videos by retrieving and composing short-horizon human clips. This approach facilitates effective policy training without the need for paired data. RHyME successfully imitates a range of cross-embodiment demonstrators, both in simulation and with a real human hand, achieving over 50% increase in task success compared to previous methods. We release our code and datasets at https://portal-cornell.github.io/rhyme/.
format Preprint
id arxiv_https___arxiv_org_abs_2409_06615
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle One-Shot Imitation under Mismatched Execution
Kedia, Kushal
Dan, Prithwish
Chao, Angela
Pace, Maximus Adrian
Choudhury, Sanjiban
Robotics
Artificial Intelligence
Machine Learning
Human demonstrations as prompts are a powerful way to program robots to do long-horizon manipulation tasks. However, translating these demonstrations into robot-executable actions presents significant challenges due to execution mismatches in movement styles and physical capabilities. Existing methods for human-robot translation either depend on paired data, which is infeasible to scale, or rely heavily on frame-level visual similarities that often break down in practice. To address these challenges, we propose RHyME, a novel framework that automatically pairs human and robot trajectories using sequence-level optimal transport cost functions. Given long-horizon robot demonstrations, RHyME synthesizes semantically equivalent human videos by retrieving and composing short-horizon human clips. This approach facilitates effective policy training without the need for paired data. RHyME successfully imitates a range of cross-embodiment demonstrators, both in simulation and with a real human hand, achieving over 50% increase in task success compared to previous methods. We release our code and datasets at https://portal-cornell.github.io/rhyme/.
title One-Shot Imitation under Mismatched Execution
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2409.06615