IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Viel, Stefano, Viano, Luca, Cevher, Volkan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910975684247552
author Viel, Stefano
Viano, Luca
Cevher, Volkan
author_facet Viel, Stefano
Viano, Luca
Cevher, Volkan
contents This paper introduces the SOAR framework for imitation learning. SOAR is an algorithmic template that learns a policy from expert demonstrations with a primal dual style algorithm that alternates cost and policy updates. Within the policy updates, the SOAR framework uses an actor critic method with multiple critics to estimate the critic uncertainty and build an optimistic critic fundamental to drive exploration. When instantiated in the tabular setting, we get a provable algorithm with guarantees that matches the best known results in $ε$. Practically, the SOAR template is shown to boost consistently the performance of imitation learning algorithms based on Soft Actor Critic such as f-IRL, ML-IRL and CSIL in several MuJoCo environments. Overall, thanks to SOAR, the required number of episodes to achieve the same performance is reduced by half.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19859
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
Viel, Stefano
Viano, Luca
Cevher, Volkan
Machine Learning
This paper introduces the SOAR framework for imitation learning. SOAR is an algorithmic template that learns a policy from expert demonstrations with a primal dual style algorithm that alternates cost and policy updates. Within the policy updates, the SOAR framework uses an actor critic method with multiple critics to estimate the critic uncertainty and build an optimistic critic fundamental to drive exploration. When instantiated in the tabular setting, we get a provable algorithm with guarantees that matches the best known results in $ε$. Practically, the SOAR template is shown to boost consistently the performance of imitation learning algorithms based on Soft Actor Critic such as f-IRL, ML-IRL and CSIL in several MuJoCo environments. Overall, thanks to SOAR, the required number of episodes to achieve the same performance is reduced by half.
title IL-SOAR : Imitation Learning with Soft Optimistic Actor cRitic
topic Machine Learning
url https://arxiv.org/abs/2502.19859