Adaptive Guidance with Reinforcement Meta-Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gaudet, Brian, Linares, Richard
Format: Preprint
Published: 2019
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912161505214464
author Gaudet, Brian
Linares, Richard
author_facet Gaudet, Brian
Linares, Richard
contents This paper proposes a novel adaptive guidance system developed using reinforcement meta-learning with a recurrent policy and value function approximator. The use of recurrent network layers allows the deployed policy to adapt real time to environmental forces acting on the agent. We compare the performance of the DR/DV guidance law, an RL agent with a non-recurrent policy, and an RL agent with a recurrent policy in four difficult tasks with unknown but highly variable dynamics. These tasks include a safe Mars landing with random engine failure and a landing on an asteroid with unknown environmental dynamics. We also demonstrate the ability of a recurrent policy to navigate using only Doppler radar altimeter returns, thus integrating guidance and navigation.
format Preprint
id arxiv_https___arxiv_org_abs_1901_04473
institution arXiv
publishDate 2019
record_format arxiv
spellingShingle Adaptive Guidance with Reinforcement Meta-Learning
Gaudet, Brian
Linares, Richard
Systems and Control
Robotics
This paper proposes a novel adaptive guidance system developed using reinforcement meta-learning with a recurrent policy and value function approximator. The use of recurrent network layers allows the deployed policy to adapt real time to environmental forces acting on the agent. We compare the performance of the DR/DV guidance law, an RL agent with a non-recurrent policy, and an RL agent with a recurrent policy in four difficult tasks with unknown but highly variable dynamics. These tasks include a safe Mars landing with random engine failure and a landing on an asteroid with unknown environmental dynamics. We also demonstrate the ability of a recurrent policy to navigate using only Doppler radar altimeter returns, thus integrating guidance and navigation.
title Adaptive Guidance with Reinforcement Meta-Learning
topic Systems and Control
Robotics
url https://arxiv.org/abs/1901.04473