Transfer in Reinforcement Learning via Regret Bounds for Learning Agents

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tuynman, Adrienne, Ortner, Ronald
Format: Preprint
Veröffentlicht: 2022
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917076881375232
author Tuynman, Adrienne
Ortner, Ronald
author_facet Tuynman, Adrienne
Ortner, Ronald
contents We present an approach for the quantification of the usefulness of transfer in reinforcement learning via regret bounds for a multi-agent setting. Considering a number of $\aleph$ agents operating in the same Markov decision process, however possibly with different reward functions, we consider the regret each agent suffers with respect to an optimal policy maximizing her average reward. We show that when the agents share their observations the total regret of all agents is smaller by a factor of $\sqrt{\aleph}$ compared to the case when each agent has to rely on the information collected by herself. This result demonstrates how considering the regret in multi-agent settings can provide theoretical bounds on the benefit of sharing observations in transfer learning.
format Preprint
id arxiv_https___arxiv_org_abs_2202_01182
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Transfer in Reinforcement Learning via Regret Bounds for Learning Agents
Tuynman, Adrienne
Ortner, Ronald
Machine Learning
Multiagent Systems
We present an approach for the quantification of the usefulness of transfer in reinforcement learning via regret bounds for a multi-agent setting. Considering a number of $\aleph$ agents operating in the same Markov decision process, however possibly with different reward functions, we consider the regret each agent suffers with respect to an optimal policy maximizing her average reward. We show that when the agents share their observations the total regret of all agents is smaller by a factor of $\sqrt{\aleph}$ compared to the case when each agent has to rely on the information collected by herself. This result demonstrates how considering the regret in multi-agent settings can provide theoretical bounds on the benefit of sharing observations in transfer learning.
title Transfer in Reinforcement Learning via Regret Bounds for Learning Agents
topic Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2202.01182