Imitation from Diverse Behaviors: Wasserstein Quality Diversity Imitation Learning with Single-Step Archive Exploration

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yu, Xingrui, Wan, Zhenglin, Bossens, David Mark, Lyu, Yueming, Guo, Qing, Tsang, Ivor W.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912308653981696
author Yu, Xingrui
Wan, Zhenglin
Bossens, David Mark
Lyu, Yueming
Guo, Qing
Tsang, Ivor W.
author_facet Yu, Xingrui
Wan, Zhenglin
Bossens, David Mark
Lyu, Yueming
Guo, Qing
Tsang, Ivor W.
contents Learning diverse and high-performance behaviors from a limited set of demonstrations is a grand challenge. Traditional imitation learning methods usually fail in this task because most of them are designed to learn one specific behavior even with multiple demonstrations. Therefore, novel techniques for \textit{quality diversity imitation learning}, which bridges the quality diversity optimization and imitation learning methods, are needed to solve the above challenge. This work introduces Wasserstein Quality Diversity Imitation Learning (WQDIL), which 1) improves the stability of imitation learning in the quality diversity setting with latent adversarial training based on a Wasserstein Auto-Encoder (WAE), and 2) mitigates a behavior-overfitting issue using a measure-conditioned reward function with a single-step archive exploration bonus. Empirically, our method significantly outperforms state-of-the-art IL methods, achieving near-expert or beyond-expert QD performance on the challenging continuous control tasks derived from MuJoCo environments.
format Preprint
id arxiv_https___arxiv_org_abs_2411_06965
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Imitation from Diverse Behaviors: Wasserstein Quality Diversity Imitation Learning with Single-Step Archive Exploration
Yu, Xingrui
Wan, Zhenglin
Bossens, David Mark
Lyu, Yueming
Guo, Qing
Tsang, Ivor W.
Machine Learning
Artificial Intelligence
Learning diverse and high-performance behaviors from a limited set of demonstrations is a grand challenge. Traditional imitation learning methods usually fail in this task because most of them are designed to learn one specific behavior even with multiple demonstrations. Therefore, novel techniques for \textit{quality diversity imitation learning}, which bridges the quality diversity optimization and imitation learning methods, are needed to solve the above challenge. This work introduces Wasserstein Quality Diversity Imitation Learning (WQDIL), which 1) improves the stability of imitation learning in the quality diversity setting with latent adversarial training based on a Wasserstein Auto-Encoder (WAE), and 2) mitigates a behavior-overfitting issue using a measure-conditioned reward function with a single-step archive exploration bonus. Empirically, our method significantly outperforms state-of-the-art IL methods, achieving near-expert or beyond-expert QD performance on the challenging continuous control tasks derived from MuJoCo environments.
title Imitation from Diverse Behaviors: Wasserstein Quality Diversity Imitation Learning with Single-Step Archive Exploration
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2411.06965