Trust Region Reward Optimization and Proximal Inverse Reward Optimization Algorithm

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Yang, Zou, Menglin, Zhang, Jiaqi, Zhang, Yitan, Yang, Junyi, Gendron, Gael, Zhang, Libo, Liu, Jiamou, Witbrock, Michael J.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912645011996672
author Chen, Yang
Zou, Menglin
Zhang, Jiaqi
Zhang, Yitan
Yang, Junyi
Gendron, Gael
Zhang, Libo
Liu, Jiamou
Witbrock, Michael J.
author_facet Chen, Yang
Zou, Menglin
Zhang, Jiaqi
Zhang, Yitan
Yang, Junyi
Gendron, Gael
Zhang, Libo
Liu, Jiamou
Witbrock, Michael J.
contents Inverse Reinforcement Learning (IRL) learns a reward function to explain expert demonstrations. Modern IRL methods often use the adversarial (minimax) formulation that alternates between reward and policy optimization, which often lead to unstable training. Recent non-adversarial IRL approaches improve stability by jointly learning reward and policy via energy-based formulations but lack formal guarantees. This work bridges this gap. We first present a unified view showing canonical non-adversarial methods explicitly or implicitly maximize the likelihood of expert behavior, which is equivalent to minimizing the expected return gap. This insight leads to our main contribution: Trust Region Reward Optimization (TRRO), a framework that guarantees monotonic improvement in this likelihood via a Minorization-Maximization process. We instantiate TRRO into Proximal Inverse Reward Optimization (PIRO), a practical and stable IRL algorithm. Theoretically, TRRO provides the IRL counterpart to the stability guarantees of Trust Region Policy Optimization (TRPO) in forward RL. Empirically, PIRO matches or surpasses state-of-the-art baselines in reward recovery, policy imitation with high sample efficiency on MuJoCo and Gym-Robotics benchmarks and a real-world animal behavior modeling task.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23135
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Trust Region Reward Optimization and Proximal Inverse Reward Optimization Algorithm
Chen, Yang
Zou, Menglin
Zhang, Jiaqi
Zhang, Yitan
Yang, Junyi
Gendron, Gael
Zhang, Libo
Liu, Jiamou
Witbrock, Michael J.
Machine Learning
Artificial Intelligence
Inverse Reinforcement Learning (IRL) learns a reward function to explain expert demonstrations. Modern IRL methods often use the adversarial (minimax) formulation that alternates between reward and policy optimization, which often lead to unstable training. Recent non-adversarial IRL approaches improve stability by jointly learning reward and policy via energy-based formulations but lack formal guarantees. This work bridges this gap. We first present a unified view showing canonical non-adversarial methods explicitly or implicitly maximize the likelihood of expert behavior, which is equivalent to minimizing the expected return gap. This insight leads to our main contribution: Trust Region Reward Optimization (TRRO), a framework that guarantees monotonic improvement in this likelihood via a Minorization-Maximization process. We instantiate TRRO into Proximal Inverse Reward Optimization (PIRO), a practical and stable IRL algorithm. Theoretically, TRRO provides the IRL counterpart to the stability guarantees of Trust Region Policy Optimization (TRPO) in forward RL. Empirically, PIRO matches or surpasses state-of-the-art baselines in reward recovery, policy imitation with high sample efficiency on MuJoCo and Gym-Robotics benchmarks and a real-world animal behavior modeling task.
title Trust Region Reward Optimization and Proximal Inverse Reward Optimization Algorithm
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.23135