Rethinking Optimal Transport in Offline Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Asadulaev, Arip, Korst, Rostislav, Korotin, Alexander, Egiazarian, Vage, Filchenkov, Andrey, Burnaev, Evgeny
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909354044686336
author Asadulaev, Arip
Korst, Rostislav
Korotin, Alexander
Egiazarian, Vage
Filchenkov, Andrey
Burnaev, Evgeny
author_facet Asadulaev, Arip
Korst, Rostislav
Korotin, Alexander
Egiazarian, Vage
Filchenkov, Andrey
Burnaev, Evgeny
contents We propose a novel algorithm for offline reinforcement learning using optimal transport. Typically, in offline reinforcement learning, the data is provided by various experts and some of them can be sub-optimal. To extract an efficient policy, it is necessary to \emph{stitch} the best behaviors from the dataset. To address this problem, we rethink offline reinforcement learning as an optimal transportation problem. And based on this, we present an algorithm that aims to find a policy that maps states to a \emph{partial} distribution of the best expert actions for each given state. We evaluate the performance of our algorithm on continuous control problems from the D4RL suite and demonstrate improvements over existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14069
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rethinking Optimal Transport in Offline Reinforcement Learning
Asadulaev, Arip
Korst, Rostislav
Korotin, Alexander
Egiazarian, Vage
Filchenkov, Andrey
Burnaev, Evgeny
Machine Learning
We propose a novel algorithm for offline reinforcement learning using optimal transport. Typically, in offline reinforcement learning, the data is provided by various experts and some of them can be sub-optimal. To extract an efficient policy, it is necessary to \emph{stitch} the best behaviors from the dataset. To address this problem, we rethink offline reinforcement learning as an optimal transportation problem. And based on this, we present an algorithm that aims to find a policy that maps states to a \emph{partial} distribution of the best expert actions for each given state. We evaluate the performance of our algorithm on continuous control problems from the D4RL suite and demonstrate improvements over existing methods.
title Rethinking Optimal Transport in Offline Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2410.14069