Enhancing Inverse Reinforcement Learning through Encoding Dynamic Information in Reward Shaping

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhan, Simon Sinong, Wang, Philip, Wu, Qingyuan, Wang, Yixuan, Jiao, Ruochen, Huang, Chao, Zhu, Qi
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918331506753536
author Zhan, Simon Sinong
Wang, Philip
Wu, Qingyuan
Wang, Yixuan
Jiao, Ruochen
Huang, Chao
Zhu, Qi
author_facet Zhan, Simon Sinong
Wang, Philip
Wu, Qingyuan
Wang, Yixuan
Jiao, Ruochen
Huang, Chao
Zhu, Qi
contents In this paper, we aim to tackle the limitation of the Adversarial Inverse Reinforcement Learning (AIRL) method in stochastic environments where theoretical results cannot hold and performance is degraded. To address this issue, we propose a novel method which infuses the dynamics information into the reward shaping with the theoretical guarantee for the induced optimal policy in the stochastic environments. Incorporating our novel model-enhanced rewards, we present a novel Model-Enhanced AIRL framework, which integrates transition model estimation directly into reward shaping. Furthermore, we provide a comprehensive theoretical analysis of the reward error bound and performance difference bound for our method. The experimental results in MuJoCo benchmarks show that our method can achieve superior performance in stochastic environments and competitive performance in deterministic environments, with significant improvement in sample efficiency, compared to existing baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03847
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Inverse Reinforcement Learning through Encoding Dynamic Information in Reward Shaping
Zhan, Simon Sinong
Wang, Philip
Wu, Qingyuan
Wang, Yixuan
Jiao, Ruochen
Huang, Chao
Zhu, Qi
Machine Learning
Artificial Intelligence
In this paper, we aim to tackle the limitation of the Adversarial Inverse Reinforcement Learning (AIRL) method in stochastic environments where theoretical results cannot hold and performance is degraded. To address this issue, we propose a novel method which infuses the dynamics information into the reward shaping with the theoretical guarantee for the induced optimal policy in the stochastic environments. Incorporating our novel model-enhanced rewards, we present a novel Model-Enhanced AIRL framework, which integrates transition model estimation directly into reward shaping. Furthermore, we provide a comprehensive theoretical analysis of the reward error bound and performance difference bound for our method. The experimental results in MuJoCo benchmarks show that our method can achieve superior performance in stochastic environments and competitive performance in deterministic environments, with significant improvement in sample efficiency, compared to existing baselines.
title Enhancing Inverse Reinforcement Learning through Encoding Dynamic Information in Reward Shaping
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.03847