Hybrid-AIRL: Enhancing Inverse Reinforcement Learning with Supervised Expert Guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Silue, Bram, Amaya-Corredor, Santiago, Mannion, Patrick, Willem, Lander, Libin, Pieter
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914496820281344
author Silue, Bram
Amaya-Corredor, Santiago
Mannion, Patrick
Willem, Lander
Libin, Pieter
author_facet Silue, Bram
Amaya-Corredor, Santiago
Mannion, Patrick
Willem, Lander
Libin, Pieter
contents Adversarial Inverse Reinforcement Learning (AIRL) has shown promise in addressing the sparse reward problem in reinforcement learning (RL) by inferring dense reward functions from expert demonstrations. However, its performance in highly complex, imperfect-information settings remains largely unexplored. To explore this gap, we evaluate AIRL in the context of Heads-Up Limit Hold'em (HULHE) poker, a domain characterized by sparse, delayed rewards and significant uncertainty. In this setting, we find that AIRL struggles to infer a sufficiently informative reward function. To overcome this limitation, we contribute Hybrid-AIRL (H-AIRL), an extension that enhances reward inference and policy learning by incorporating a supervised loss derived from expert data and a stochastic regularization mechanism. We evaluate H-AIRL on a carefully selected set of Gymnasium benchmarks and the HULHE poker setting. Additionally, we analyze the learned reward function through visualization to gain deeper insights into the learning process. Our experimental results show that H-AIRL achieves higher sample efficiency and more stable learning compared to AIRL. This highlights the benefits of incorporating supervised signals into inverse RL and establishes H-AIRL as a promising framework for tackling challenging, real-world settings.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21356
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hybrid-AIRL: Enhancing Inverse Reinforcement Learning with Supervised Expert Guidance
Silue, Bram
Amaya-Corredor, Santiago
Mannion, Patrick
Willem, Lander
Libin, Pieter
Machine Learning
Artificial Intelligence
I.2.6; I.2.8; I.2.1
Adversarial Inverse Reinforcement Learning (AIRL) has shown promise in addressing the sparse reward problem in reinforcement learning (RL) by inferring dense reward functions from expert demonstrations. However, its performance in highly complex, imperfect-information settings remains largely unexplored. To explore this gap, we evaluate AIRL in the context of Heads-Up Limit Hold'em (HULHE) poker, a domain characterized by sparse, delayed rewards and significant uncertainty. In this setting, we find that AIRL struggles to infer a sufficiently informative reward function. To overcome this limitation, we contribute Hybrid-AIRL (H-AIRL), an extension that enhances reward inference and policy learning by incorporating a supervised loss derived from expert data and a stochastic regularization mechanism. We evaluate H-AIRL on a carefully selected set of Gymnasium benchmarks and the HULHE poker setting. Additionally, we analyze the learned reward function through visualization to gain deeper insights into the learning process. Our experimental results show that H-AIRL achieves higher sample efficiency and more stable learning compared to AIRL. This highlights the benefits of incorporating supervised signals into inverse RL and establishes H-AIRL as a promising framework for tackling challenging, real-world settings.
title Hybrid-AIRL: Enhancing Inverse Reinforcement Learning with Supervised Expert Guidance
topic Machine Learning
Artificial Intelligence
I.2.6; I.2.8; I.2.1
url https://arxiv.org/abs/2511.21356