Improving Controller Generalization with Dimensionless Markov Decision Processes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Charvet, Valentin, Stein, Sebastian, Murray-Smith, Roderick
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915240943288320
author Charvet, Valentin
Stein, Sebastian
Murray-Smith, Roderick
author_facet Charvet, Valentin
Stein, Sebastian
Murray-Smith, Roderick
contents Controllers trained with Reinforcement Learning tend to be very specialized and thus generalize poorly when their testing environment differs from their training one. We propose a Model-Based approach to increase generalization where both world model and policy are trained in a dimensionless state-action space. To do so, we introduce the Dimensionless Markov Decision Process ($Π$-MDP): an extension of Contextual-MDPs in which state and action spaces are non-dimensionalized with the Buckingham-$Π$ theorem. This procedure induces policies that are equivariant with respect to changes in the context of the underlying dynamics. We provide a generic framework for this approach and apply it to a model-based policy search algorithm using Gaussian Process models. We demonstrate the applicability of our method on simulated actuated pendulum and cartpole systems, where policies trained on a single environment are robust to shifts in the distribution of the context.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10006
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Controller Generalization with Dimensionless Markov Decision Processes
Charvet, Valentin
Stein, Sebastian
Murray-Smith, Roderick
Machine Learning
Controllers trained with Reinforcement Learning tend to be very specialized and thus generalize poorly when their testing environment differs from their training one. We propose a Model-Based approach to increase generalization where both world model and policy are trained in a dimensionless state-action space. To do so, we introduce the Dimensionless Markov Decision Process ($Π$-MDP): an extension of Contextual-MDPs in which state and action spaces are non-dimensionalized with the Buckingham-$Π$ theorem. This procedure induces policies that are equivariant with respect to changes in the context of the underlying dynamics. We provide a generic framework for this approach and apply it to a model-based policy search algorithm using Gaussian Process models. We demonstrate the applicability of our method on simulated actuated pendulum and cartpole systems, where policies trained on a single environment are robust to shifts in the distribution of the context.
title Improving Controller Generalization with Dimensionless Markov Decision Processes
topic Machine Learning
url https://arxiv.org/abs/2504.10006