Generalizing Behavior via Inverse Reinforcement Learning with Closed-Form Reward Centroids

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lazzati, Filippo, Metelli, Alberto Maria
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912587778621440
author Lazzati, Filippo
Metelli, Alberto Maria
author_facet Lazzati, Filippo
Metelli, Alberto Maria
contents We study the problem of generalizing an expert agent's behavior, provided through demonstrations, to new environments and/or additional constraints. Inverse Reinforcement Learning (IRL) offers a promising solution by seeking to recover the expert's underlying reward function, which, if used for planning in the new settings, would reproduce the desired behavior. However, IRL is inherently ill-posed: multiple reward functions, forming the so-called feasible set, can explain the same observed behavior. Since these rewards may induce different policies in the new setting, in the absence of additional information, a decision criterion is needed to select which policy to deploy. In this paper, we propose a novel, principled criterion that selects the "average" policy among those induced by the rewards in a certain bounded subset of the feasible set. Remarkably, we show that this policy can be obtained by planning with the reward centroid of that subset, for which we derive a closed-form expression. We then present a provably efficient algorithm for estimating this centroid using an offline dataset of expert demonstrations only. Finally, we conduct numerical simulations that illustrate the relationship between the expert's behavior and the behavior produced by our method.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12010
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generalizing Behavior via Inverse Reinforcement Learning with Closed-Form Reward Centroids
Lazzati, Filippo
Metelli, Alberto Maria
Machine Learning
Artificial Intelligence
We study the problem of generalizing an expert agent's behavior, provided through demonstrations, to new environments and/or additional constraints. Inverse Reinforcement Learning (IRL) offers a promising solution by seeking to recover the expert's underlying reward function, which, if used for planning in the new settings, would reproduce the desired behavior. However, IRL is inherently ill-posed: multiple reward functions, forming the so-called feasible set, can explain the same observed behavior. Since these rewards may induce different policies in the new setting, in the absence of additional information, a decision criterion is needed to select which policy to deploy. In this paper, we propose a novel, principled criterion that selects the "average" policy among those induced by the rewards in a certain bounded subset of the feasible set. Remarkably, we show that this policy can be obtained by planning with the reward centroid of that subset, for which we derive a closed-form expression. We then present a provably efficient algorithm for estimating this centroid using an offline dataset of expert demonstrations only. Finally, we conduct numerical simulations that illustrate the relationship between the expert's behavior and the behavior produced by our method.
title Generalizing Behavior via Inverse Reinforcement Learning with Closed-Form Reward Centroids
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.12010