Impact of Markov Decision Process Design on Sim-to-Real Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866910050216312832 |
|---|---|
| author | Krau, Tatjana Mandlmaier, Jorge Damm, Tobias Heieck, Frieder |
| author_facet | Krau, Tatjana Mandlmaier, Jorge Damm, Tobias Heieck, Frieder |
| contents | Reinforcement Learning (RL) has demonstrated strong potential for industrial process control, yet policies trained in simulation often suffer from a significant sim-to-real gap when deployed on physical hardware. This work systematically analyzes how core Markov Decision Process (MDP) design choices -- state composition, target inclusion, reward formulation, termination criteria, and environment dynamics models -- affect this transfer. Using a color mixing task, we evaluate different MDP configurations and mixing dynamics across simulation and real-world experiments. We validate our findings on physical hardware, demonstrating that physics-based dynamics models achieve up to 50% real-world success under strict precision constraints where simplified models fail entirely. Our results provide practical MDP design guidelines for deploying RL in industrial process control. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_09427 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Impact of Markov Decision Process Design on Sim-to-Real Reinforcement Learning Krau, Tatjana Mandlmaier, Jorge Damm, Tobias Heieck, Frieder Machine Learning Reinforcement Learning (RL) has demonstrated strong potential for industrial process control, yet policies trained in simulation often suffer from a significant sim-to-real gap when deployed on physical hardware. This work systematically analyzes how core Markov Decision Process (MDP) design choices -- state composition, target inclusion, reward formulation, termination criteria, and environment dynamics models -- affect this transfer. Using a color mixing task, we evaluate different MDP configurations and mixing dynamics across simulation and real-world experiments. We validate our findings on physical hardware, demonstrating that physics-based dynamics models achieve up to 50% real-world success under strict precision constraints where simplified models fail entirely. Our results provide practical MDP design guidelines for deploying RL in industrial process control. |
| title | Impact of Markov Decision Process Design on Sim-to-Real Reinforcement Learning |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2603.09427 |