Guided Reinforcement Learning for Robust Multi-Contact Loco-Manipulation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sleiman, Jean-Pierre, Mittal, Mayank, Hutter, Marco
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912075524079616
author Sleiman, Jean-Pierre
Mittal, Mayank
Hutter, Marco
author_facet Sleiman, Jean-Pierre
Mittal, Mayank
Hutter, Marco
contents Reinforcement learning (RL) often necessitates a meticulous Markov Decision Process (MDP) design tailored to each task. This work aims to address this challenge by proposing a systematic approach to behavior synthesis and control for multi-contact loco-manipulation tasks, such as navigating spring-loaded doors and manipulating heavy dishwashers. We define a task-independent MDP to train RL policies using only a single demonstration per task generated from a model-based trajectory optimizer. Our approach incorporates an adaptive phase dynamics formulation to robustly track the demonstrations while accommodating dynamic uncertainties and external disturbances. We compare our method against prior motion imitation RL works and show that the learned policies achieve higher success rates across all considered tasks. These policies learn recovery maneuvers that are not present in the demonstration, such as re-grasping objects during execution or dealing with slippages. Finally, we successfully transfer the policies to a real robot, demonstrating the practical viability of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13817
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Guided Reinforcement Learning for Robust Multi-Contact Loco-Manipulation
Sleiman, Jean-Pierre
Mittal, Mayank
Hutter, Marco
Robotics
Artificial Intelligence
Reinforcement learning (RL) often necessitates a meticulous Markov Decision Process (MDP) design tailored to each task. This work aims to address this challenge by proposing a systematic approach to behavior synthesis and control for multi-contact loco-manipulation tasks, such as navigating spring-loaded doors and manipulating heavy dishwashers. We define a task-independent MDP to train RL policies using only a single demonstration per task generated from a model-based trajectory optimizer. Our approach incorporates an adaptive phase dynamics formulation to robustly track the demonstrations while accommodating dynamic uncertainties and external disturbances. We compare our method against prior motion imitation RL works and show that the learned policies achieve higher success rates across all considered tasks. These policies learn recovery maneuvers that are not present in the demonstration, such as re-grasping objects during execution or dealing with slippages. Finally, we successfully transfer the policies to a real robot, demonstrating the practical viability of our approach.
title Guided Reinforcement Learning for Robust Multi-Contact Loco-Manipulation
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2410.13817