On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Freihaut, Till, Ramponi, Giorgia
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912728332894208
author Freihaut, Till
Ramponi, Giorgia
author_facet Freihaut, Till
Ramponi, Giorgia
contents Multi-agent Inverse Reinforcement Learning (MAIRL) aims to recover agent reward functions from expert demonstrations. We characterize the feasible reward set in Markov games, identifying all reward functions that rationalize a given equilibrium. However, equilibrium-based observations are often ambiguous: a single Nash equilibrium can correspond to many reward structures, potentially changing the game's nature in multi-agent systems. We address this by introducing entropy-regularized Markov games, which yield a unique equilibrium while preserving strategic incentives. For this setting, we provide a sample complexity analysis detailing how errors affect learned policy performance. Our work establishes theoretical foundations and practical insights for MAIRL.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15046
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning
Freihaut, Till
Ramponi, Giorgia
Machine Learning
Multi-agent Inverse Reinforcement Learning (MAIRL) aims to recover agent reward functions from expert demonstrations. We characterize the feasible reward set in Markov games, identifying all reward functions that rationalize a given equilibrium. However, equilibrium-based observations are often ambiguous: a single Nash equilibrium can correspond to many reward structures, potentially changing the game's nature in multi-agent systems. We address this by introducing entropy-regularized Markov games, which yield a unique equilibrium while preserving strategic incentives. For this setting, we provide a sample complexity analysis detailing how errors affect learned policy performance. Our work establishes theoretical foundations and practical insights for MAIRL.
title On Feasible Rewards in Multi-Agent Inverse Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2411.15046