Offline Learning of Controllable Diverse Behaviors

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Petitbois, Mathieu, Portelas, Rémy, Lamprier, Sylvain, Denoyer, Ludovic
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910918787465216
author Petitbois, Mathieu
Portelas, Rémy
Lamprier, Sylvain
Denoyer, Ludovic
author_facet Petitbois, Mathieu
Portelas, Rémy
Lamprier, Sylvain
Denoyer, Ludovic
contents Imitation Learning (IL) techniques aim to replicate human behaviors in specific tasks. While IL has gained prominence due to its effectiveness and efficiency, traditional methods often focus on datasets collected from experts to produce a single efficient policy. Recently, extensions have been proposed to handle datasets of diverse behaviors by mainly focusing on learning transition-level diverse policies or on performing entropy maximization at the trajectory level. While these methods may lead to diverse behaviors, they may not be sufficient to reproduce the actual diversity of demonstrations or to allow controlled trajectory generation. To overcome these drawbacks, we propose a different method based on two key features: a) Temporal Consistency that ensures consistent behaviors across entire episodes and not just at the transition level as well as b) Controllability obtained by constructing a latent space of behaviors that allows users to selectively activate specific behaviors based on their requirements. We compare our approach to state-of-the-art methods over a diverse set of tasks and environments. Project page: https://mathieu-petitbois.github.io/projects/swr/
format Preprint
id arxiv_https___arxiv_org_abs_2504_18160
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Offline Learning of Controllable Diverse Behaviors
Petitbois, Mathieu
Portelas, Rémy
Lamprier, Sylvain
Denoyer, Ludovic
Machine Learning
Artificial Intelligence
Robotics
Imitation Learning (IL) techniques aim to replicate human behaviors in specific tasks. While IL has gained prominence due to its effectiveness and efficiency, traditional methods often focus on datasets collected from experts to produce a single efficient policy. Recently, extensions have been proposed to handle datasets of diverse behaviors by mainly focusing on learning transition-level diverse policies or on performing entropy maximization at the trajectory level. While these methods may lead to diverse behaviors, they may not be sufficient to reproduce the actual diversity of demonstrations or to allow controlled trajectory generation. To overcome these drawbacks, we propose a different method based on two key features: a) Temporal Consistency that ensures consistent behaviors across entire episodes and not just at the transition level as well as b) Controllability obtained by constructing a latent space of behaviors that allows users to selectively activate specific behaviors based on their requirements. We compare our approach to state-of-the-art methods over a diverse set of tasks and environments. Project page: https://mathieu-petitbois.github.io/projects/swr/
title Offline Learning of Controllable Diverse Behaviors
topic Machine Learning
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2504.18160