Zero-Shot Offline Imitation Learning via Optimal Transport

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rupf, Thomas, Bagatella, Marco, Gürtler, Nico, Frey, Jonas, Martius, Georg
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915338526916608
author Rupf, Thomas
Bagatella, Marco
Gürtler, Nico
Frey, Jonas
Martius, Georg
author_facet Rupf, Thomas
Bagatella, Marco
Gürtler, Nico
Frey, Jonas
Martius, Georg
contents Zero-shot imitation learning algorithms hold the promise of reproducing unseen behavior from as little as a single demonstration at test time. Existing practical approaches view the expert demonstration as a sequence of goals, enabling imitation with a high-level goal selector, and a low-level goal-conditioned policy. However, this framework can suffer from myopic behavior: the agent's immediate actions towards achieving individual goals may undermine long-term objectives. We introduce a novel method that mitigates this issue by directly optimizing the occupancy matching objective that is intrinsic to imitation learning. We propose to lift a goal-conditioned value function to a distance between occupancies, which are in turn approximated via a learned world model. The resulting method can learn from offline, suboptimal data, and is capable of non-myopic, zero-shot imitation, as we demonstrate in complex, continuous benchmarks. The code is available at https://github.com/martius-lab/zilot.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08751
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Zero-Shot Offline Imitation Learning via Optimal Transport
Rupf, Thomas
Bagatella, Marco
Gürtler, Nico
Frey, Jonas
Martius, Georg
Machine Learning
Zero-shot imitation learning algorithms hold the promise of reproducing unseen behavior from as little as a single demonstration at test time. Existing practical approaches view the expert demonstration as a sequence of goals, enabling imitation with a high-level goal selector, and a low-level goal-conditioned policy. However, this framework can suffer from myopic behavior: the agent's immediate actions towards achieving individual goals may undermine long-term objectives. We introduce a novel method that mitigates this issue by directly optimizing the occupancy matching objective that is intrinsic to imitation learning. We propose to lift a goal-conditioned value function to a distance between occupancies, which are in turn approximated via a learned world model. The resulting method can learn from offline, suboptimal data, and is capable of non-myopic, zero-shot imitation, as we demonstrate in complex, continuous benchmarks. The code is available at https://github.com/martius-lab/zilot.
title Zero-Shot Offline Imitation Learning via Optimal Transport
topic Machine Learning
url https://arxiv.org/abs/2410.08751