Impact of Markov Decision Process Design on Sim-to-Real Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Krau, Tatjana, Mandlmaier, Jorge, Damm, Tobias, Heieck, Frieder
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910050216312832
author Krau, Tatjana
Mandlmaier, Jorge
Damm, Tobias
Heieck, Frieder
author_facet Krau, Tatjana
Mandlmaier, Jorge
Damm, Tobias
Heieck, Frieder
contents Reinforcement Learning (RL) has demonstrated strong potential for industrial process control, yet policies trained in simulation often suffer from a significant sim-to-real gap when deployed on physical hardware. This work systematically analyzes how core Markov Decision Process (MDP) design choices -- state composition, target inclusion, reward formulation, termination criteria, and environment dynamics models -- affect this transfer. Using a color mixing task, we evaluate different MDP configurations and mixing dynamics across simulation and real-world experiments. We validate our findings on physical hardware, demonstrating that physics-based dynamics models achieve up to 50% real-world success under strict precision constraints where simplified models fail entirely. Our results provide practical MDP design guidelines for deploying RL in industrial process control.
format Preprint
id arxiv_https___arxiv_org_abs_2603_09427
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Impact of Markov Decision Process Design on Sim-to-Real Reinforcement Learning
Krau, Tatjana
Mandlmaier, Jorge
Damm, Tobias
Heieck, Frieder
Machine Learning
Reinforcement Learning (RL) has demonstrated strong potential for industrial process control, yet policies trained in simulation often suffer from a significant sim-to-real gap when deployed on physical hardware. This work systematically analyzes how core Markov Decision Process (MDP) design choices -- state composition, target inclusion, reward formulation, termination criteria, and environment dynamics models -- affect this transfer. Using a color mixing task, we evaluate different MDP configurations and mixing dynamics across simulation and real-world experiments. We validate our findings on physical hardware, demonstrating that physics-based dynamics models achieve up to 50% real-world success under strict precision constraints where simplified models fail entirely. Our results provide practical MDP design guidelines for deploying RL in industrial process control.
title Impact of Markov Decision Process Design on Sim-to-Real Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2603.09427