Model Predictive Adversarial Imitation Learning for Planning from Observation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Tyler, Bao, Yanda, Mehta, Bhaumik, Guo, Gabriel, Vishwakarma, Anubhav, Kang, Emily, Jung, Sanghun, Scalise, Rosario, Zhou, Jason, Xu, Bryan, Boots, Byron
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908859523661824
author Han, Tyler
Bao, Yanda
Mehta, Bhaumik
Guo, Gabriel
Vishwakarma, Anubhav
Kang, Emily
Jung, Sanghun
Scalise, Rosario
Zhou, Jason
Xu, Bryan
Boots, Byron
author_facet Han, Tyler
Bao, Yanda
Mehta, Bhaumik
Guo, Gabriel
Vishwakarma, Anubhav
Kang, Emily
Jung, Sanghun
Scalise, Rosario
Zhou, Jason
Xu, Bryan
Boots, Byron
contents Human demonstration data is often ambiguous and incomplete, motivating imitation learning approaches that also exhibit reliable planning behavior. A common paradigm to perform planning-from-demonstration involves learning a reward function via Inverse Reinforcement Learning (IRL) then deploying this reward via Model Predictive Control (MPC). Towards unifying these methods, we derive a replacement of the policy in IRL with a planning-based agent. With connections to Adversarial Imitation Learning, this formulation enables end-to-end interactive learning of planners from observation-only demonstrations. In addition to benefits in interpretability, complexity, and safety, we study and observe significant improvements on sample efficiency, out-of-distribution generalization, and robustness. The study includes evaluations in both simulated control benchmarks and real-world navigation experiments using few-to-single observation-only demonstrations.
format Preprint
id arxiv_https___arxiv_org_abs_2507_21533
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Model Predictive Adversarial Imitation Learning for Planning from Observation
Han, Tyler
Bao, Yanda
Mehta, Bhaumik
Guo, Gabriel
Vishwakarma, Anubhav
Kang, Emily
Jung, Sanghun
Scalise, Rosario
Zhou, Jason
Xu, Bryan
Boots, Byron
Robotics
Artificial Intelligence
Human demonstration data is often ambiguous and incomplete, motivating imitation learning approaches that also exhibit reliable planning behavior. A common paradigm to perform planning-from-demonstration involves learning a reward function via Inverse Reinforcement Learning (IRL) then deploying this reward via Model Predictive Control (MPC). Towards unifying these methods, we derive a replacement of the policy in IRL with a planning-based agent. With connections to Adversarial Imitation Learning, this formulation enables end-to-end interactive learning of planners from observation-only demonstrations. In addition to benefits in interpretability, complexity, and safety, we study and observe significant improvements on sample efficiency, out-of-distribution generalization, and robustness. The study includes evaluations in both simulated control benchmarks and real-world navigation experiments using few-to-single observation-only demonstrations.
title Model Predictive Adversarial Imitation Learning for Planning from Observation
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2507.21533