Coordinating Planning and Tracking in Layered Control Policies via Actor-Critic Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yang, Fengjun, Matni, Nikolai
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929634632794112
author Yang, Fengjun
Matni, Nikolai
author_facet Yang, Fengjun
Matni, Nikolai
contents We propose a reinforcement learning (RL)-based algorithm to jointly train (1) a trajectory planner and (2) a tracking controller in a layered control architecture. Our algorithm arises naturally from a rewrite of the underlying optimal control problem that lends itself to an actor-critic learning approach. By explicitly learning a \textit{dual} network to coordinate the interaction between the planning and tracking layers, we demonstrate the ability to achieve an effective consensus between the two components, leading to an interpretable policy. We theoretically prove that our algorithm converges to the optimal dual network in the Linear Quadratic Regulator (LQR) setting and empirically validate its applicability to nonlinear systems through simulation experiments on a unicycle model.
format Preprint
id arxiv_https___arxiv_org_abs_2408_01639
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Coordinating Planning and Tracking in Layered Control Policies via Actor-Critic Learning
Yang, Fengjun
Matni, Nikolai
Systems and Control
Machine Learning
We propose a reinforcement learning (RL)-based algorithm to jointly train (1) a trajectory planner and (2) a tracking controller in a layered control architecture. Our algorithm arises naturally from a rewrite of the underlying optimal control problem that lends itself to an actor-critic learning approach. By explicitly learning a \textit{dual} network to coordinate the interaction between the planning and tracking layers, we demonstrate the ability to achieve an effective consensus between the two components, leading to an interpretable policy. We theoretically prove that our algorithm converges to the optimal dual network in the Linear Quadratic Regulator (LQR) setting and empirically validate its applicability to nonlinear systems through simulation experiments on a unicycle model.
title Coordinating Planning and Tracking in Layered Control Policies via Actor-Critic Learning
topic Systems and Control
Machine Learning
url https://arxiv.org/abs/2408.01639