Learning Control Policies for Variable Objectives from Offline Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Weber, Marc, Swazinna, Phillip, Hein, Daniel, Udluft, Steffen, Sterzing, Volkmar
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911746096103424
author Weber, Marc
Swazinna, Phillip
Hein, Daniel
Udluft, Steffen
Sterzing, Volkmar
author_facet Weber, Marc
Swazinna, Phillip
Hein, Daniel
Udluft, Steffen
Sterzing, Volkmar
contents Offline reinforcement learning provides a viable approach to obtain advanced control strategies for dynamical systems, in particular when direct interaction with the environment is not available. In this paper, we introduce a conceptual extension for model-based policy search methods, called variable objective policy (VOP). With this approach, policies are trained to generalize efficiently over a variety of objectives, which parameterize the reward function. We demonstrate that by altering the objectives passed as input to the policy, users gain the freedom to adjust its behavior or re-balance optimization targets at runtime, without need for collecting additional observation batches or re-training.
format Preprint
id arxiv_https___arxiv_org_abs_2308_06127
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Learning Control Policies for Variable Objectives from Offline Data
Weber, Marc
Swazinna, Phillip
Hein, Daniel
Udluft, Steffen
Sterzing, Volkmar
Machine Learning
Offline reinforcement learning provides a viable approach to obtain advanced control strategies for dynamical systems, in particular when direct interaction with the environment is not available. In this paper, we introduce a conceptual extension for model-based policy search methods, called variable objective policy (VOP). With this approach, policies are trained to generalize efficiently over a variety of objectives, which parameterize the reward function. We demonstrate that by altering the objectives passed as input to the policy, users gain the freedom to adjust its behavior or re-balance optimization targets at runtime, without need for collecting additional observation batches or re-training.
title Learning Control Policies for Variable Objectives from Offline Data
topic Machine Learning
url https://arxiv.org/abs/2308.06127