Recursive Deep Inverse Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghanem, Paul, Howell, Owen, Potter, Michael, Closas, Pau, Ramezani, Alireza, Erdogmus, Deniz, Imbiriba, Tales
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915533205536768
author Ghanem, Paul
Howell, Owen
Potter, Michael
Closas, Pau
Ramezani, Alireza
Erdogmus, Deniz
Imbiriba, Tales
author_facet Ghanem, Paul
Howell, Owen
Potter, Michael
Closas, Pau
Ramezani, Alireza
Erdogmus, Deniz
Imbiriba, Tales
contents Inferring an adversary's goals from exhibited behavior is crucial for counterplanning and non-cooperative multi-agent systems in domains like cybersecurity, military, and strategy games. Deep Inverse Reinforcement Learning (IRL) methods based on maximum entropy principles show promise in recovering adversaries' goals but are typically offline, require large batch sizes with gradient descent, and rely on first-order updates, limiting their applicability in real-time scenarios. We propose an online Recursive Deep Inverse Reinforcement Learning (RDIRL) approach to recover the cost function governing the adversary actions and goals. Specifically, we minimize an upper bound on the standard Guided Cost Learning (GCL) objective using sequential second-order Newton updates, akin to the Extended Kalman Filter (EKF), leading to a fast (in terms of convergence) learning algorithm. We demonstrate that RDIRL is able to recover cost and reward functions of expert agents in standard and adversarial benchmark tasks. Experiments on benchmark tasks show that our proposed approach outperforms several leading IRL algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13241
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Recursive Deep Inverse Reinforcement Learning
Ghanem, Paul
Howell, Owen
Potter, Michael
Closas, Pau
Ramezani, Alireza
Erdogmus, Deniz
Imbiriba, Tales
Machine Learning
Artificial Intelligence
Inferring an adversary's goals from exhibited behavior is crucial for counterplanning and non-cooperative multi-agent systems in domains like cybersecurity, military, and strategy games. Deep Inverse Reinforcement Learning (IRL) methods based on maximum entropy principles show promise in recovering adversaries' goals but are typically offline, require large batch sizes with gradient descent, and rely on first-order updates, limiting their applicability in real-time scenarios. We propose an online Recursive Deep Inverse Reinforcement Learning (RDIRL) approach to recover the cost function governing the adversary actions and goals. Specifically, we minimize an upper bound on the standard Guided Cost Learning (GCL) objective using sequential second-order Newton updates, akin to the Extended Kalman Filter (EKF), leading to a fast (in terms of convergence) learning algorithm. We demonstrate that RDIRL is able to recover cost and reward functions of expert agents in standard and adversarial benchmark tasks. Experiments on benchmark tasks show that our proposed approach outperforms several leading IRL algorithms.
title Recursive Deep Inverse Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2504.13241