Dual Action Policy for Robust Sim-to-Real Reinforcement Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Terence, Ng Wen Zheng, Jianda, Chen
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914974501175296
author Terence, Ng Wen Zheng
Jianda, Chen
author_facet Terence, Ng Wen Zheng
Jianda, Chen
contents This paper presents Dual Action Policy (DAP), a novel approach to address the dynamics mismatch inherent in the sim-to-real gap of reinforcement learning. DAP uses a single policy to predict two sets of actions: one for maximizing task rewards in simulation and another specifically for domain adaptation via reward adjustments. This decoupling makes it easier to maximize the overall reward in the source domain during training. Additionally, DAP incorporates uncertainty-based exploration during training to enhance agent robustness. Experimental results demonstrate DAP's effectiveness in bridging the sim-to-real gap, outperforming baselines on challenging tasks in simulation, and further improvement is achieved by incorporating uncertainty estimation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12250
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dual Action Policy for Robust Sim-to-Real Reinforcement Learning
Terence, Ng Wen Zheng
Jianda, Chen
Machine Learning
Artificial Intelligence
Robotics
This paper presents Dual Action Policy (DAP), a novel approach to address the dynamics mismatch inherent in the sim-to-real gap of reinforcement learning. DAP uses a single policy to predict two sets of actions: one for maximizing task rewards in simulation and another specifically for domain adaptation via reward adjustments. This decoupling makes it easier to maximize the overall reward in the source domain during training. Additionally, DAP incorporates uncertainty-based exploration during training to enhance agent robustness. Experimental results demonstrate DAP's effectiveness in bridging the sim-to-real gap, outperforming baselines on challenging tasks in simulation, and further improvement is achieved by incorporating uncertainty estimation.
title Dual Action Policy for Robust Sim-to-Real Reinforcement Learning
topic Machine Learning
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2410.12250