RecoveryChaining: Learning Local Recovery Policies for Robust Manipulation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Vats, Shivam, Jha, Devesh K., Likhachev, Maxim, Kroemer, Oliver, Romeres, Diego
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909528307531776
author Vats, Shivam
Jha, Devesh K.
Likhachev, Maxim
Kroemer, Oliver
Romeres, Diego
author_facet Vats, Shivam
Jha, Devesh K.
Likhachev, Maxim
Kroemer, Oliver
Romeres, Diego
contents Model-based planners and controllers are commonly used to solve complex manipulation problems as they can efficiently optimize diverse objectives and generalize to long horizon tasks. However, they often fail during deployment due to noisy actuation, partial observability and imperfect models. To enable a robot to recover from such failures, we propose to use hierarchical reinforcement learning to learn a recovery policy. The recovery policy is triggered when a failure is detected based on sensory observations and seeks to take the robot to a state from which it can complete the task using the nominal model-based controllers. Our approach, called RecoveryChaining, uses a hybrid action space, where the model-based controllers are provided as additional \emph{nominal} options which allows the recovery policy to decide how to recover, when to switch to a nominal controller and which controller to switch to even with \emph{sparse rewards}. We evaluate our approach in three multi-step manipulation tasks with sparse rewards, where it learns significantly more robust recovery policies than those learned by baselines. We successfully transfer recovery policies learned in simulation to a physical robot to demonstrate the feasibility of sim-to-real transfer with our method.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13979
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RecoveryChaining: Learning Local Recovery Policies for Robust Manipulation
Vats, Shivam
Jha, Devesh K.
Likhachev, Maxim
Kroemer, Oliver
Romeres, Diego
Robotics
Artificial Intelligence
Model-based planners and controllers are commonly used to solve complex manipulation problems as they can efficiently optimize diverse objectives and generalize to long horizon tasks. However, they often fail during deployment due to noisy actuation, partial observability and imperfect models. To enable a robot to recover from such failures, we propose to use hierarchical reinforcement learning to learn a recovery policy. The recovery policy is triggered when a failure is detected based on sensory observations and seeks to take the robot to a state from which it can complete the task using the nominal model-based controllers. Our approach, called RecoveryChaining, uses a hybrid action space, where the model-based controllers are provided as additional \emph{nominal} options which allows the recovery policy to decide how to recover, when to switch to a nominal controller and which controller to switch to even with \emph{sparse rewards}. We evaluate our approach in three multi-step manipulation tasks with sparse rewards, where it learns significantly more robust recovery policies than those learned by baselines. We successfully transfer recovery policies learned in simulation to a physical robot to demonstrate the feasibility of sim-to-real transfer with our method.
title RecoveryChaining: Learning Local Recovery Policies for Robust Manipulation
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2410.13979