Learning Transparent Reward Models via Unsupervised Feature Selection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Baimukashev, Daulet, Alcan, Gokhan, Luck, Kevin Sebastian, Kyrki, Ville
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908345078644736
author Baimukashev, Daulet
Alcan, Gokhan
Luck, Kevin Sebastian
Kyrki, Ville
author_facet Baimukashev, Daulet
Alcan, Gokhan
Luck, Kevin Sebastian
Kyrki, Ville
contents In complex real-world tasks such as robotic manipulation and autonomous driving, collecting expert demonstrations is often more straightforward than specifying precise learning objectives and task descriptions. Learning from expert data can be achieved through behavioral cloning or by learning a reward function, i.e., inverse reinforcement learning. The latter allows for training with additional data outside the training distribution, guided by the inferred reward function. We propose a novel approach to construct compact and transparent reward models from automatically selected state features. These inferred rewards have an explicit form and enable the learning of policies that closely match expert behavior by training standard reinforcement learning algorithms from scratch. We validate our method's performance in various robotic environments with continuous and high-dimensional state spaces. Webpage: \url{https://sites.google.com/view/transparent-reward}.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18608
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Transparent Reward Models via Unsupervised Feature Selection
Baimukashev, Daulet
Alcan, Gokhan
Luck, Kevin Sebastian
Kyrki, Ville
Robotics
In complex real-world tasks such as robotic manipulation and autonomous driving, collecting expert demonstrations is often more straightforward than specifying precise learning objectives and task descriptions. Learning from expert data can be achieved through behavioral cloning or by learning a reward function, i.e., inverse reinforcement learning. The latter allows for training with additional data outside the training distribution, guided by the inferred reward function. We propose a novel approach to construct compact and transparent reward models from automatically selected state features. These inferred rewards have an explicit form and enable the learning of policies that closely match expert behavior by training standard reinforcement learning algorithms from scratch. We validate our method's performance in various robotic environments with continuous and high-dimensional state spaces. Webpage: \url{https://sites.google.com/view/transparent-reward}.
title Learning Transparent Reward Models via Unsupervised Feature Selection
topic Robotics
url https://arxiv.org/abs/2410.18608