Temporal Logic Imitation: Learning Plan-Satisficing Motion Policies from Demonstrations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yanwei, Figueroa, Nadia, Li, Shen, Shah, Ankit, Shah, Julie
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913600225935360
author Wang, Yanwei
Figueroa, Nadia
Li, Shen
Shah, Ankit
Shah, Julie
author_facet Wang, Yanwei
Figueroa, Nadia
Li, Shen
Shah, Ankit
Shah, Julie
contents Learning from demonstration (LfD) has succeeded in tasks featuring a long time horizon. However, when the problem complexity also includes human-in-the-loop perturbations, state-of-the-art approaches do not guarantee the successful reproduction of a task. In this work, we identify the roots of this challenge as the failure of a learned continuous policy to satisfy the discrete plan implicit in the demonstration. By utilizing modes (rather than subgoals) as the discrete abstraction and motion policies with both mode invariance and goal reachability properties, we prove our learned continuous policy can simulate any discrete plan specified by a linear temporal logic (LTL) formula. Consequently, an imitator is robust to both task- and motion-level perturbations and guaranteed to achieve task success. Project page: https://yanweiw.github.io/tli/
format Preprint
id arxiv_https___arxiv_org_abs_2206_04632
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Temporal Logic Imitation: Learning Plan-Satisficing Motion Policies from Demonstrations
Wang, Yanwei
Figueroa, Nadia
Li, Shen
Shah, Ankit
Shah, Julie
Robotics
Artificial Intelligence
Formal Languages and Automata Theory
Machine Learning
Systems and Control
Learning from demonstration (LfD) has succeeded in tasks featuring a long time horizon. However, when the problem complexity also includes human-in-the-loop perturbations, state-of-the-art approaches do not guarantee the successful reproduction of a task. In this work, we identify the roots of this challenge as the failure of a learned continuous policy to satisfy the discrete plan implicit in the demonstration. By utilizing modes (rather than subgoals) as the discrete abstraction and motion policies with both mode invariance and goal reachability properties, we prove our learned continuous policy can simulate any discrete plan specified by a linear temporal logic (LTL) formula. Consequently, an imitator is robust to both task- and motion-level perturbations and guaranteed to achieve task success. Project page: https://yanweiw.github.io/tli/
title Temporal Logic Imitation: Learning Plan-Satisficing Motion Policies from Demonstrations
topic Robotics
Artificial Intelligence
Formal Languages and Automata Theory
Machine Learning
Systems and Control
url https://arxiv.org/abs/2206.04632