Goal-Conditioned Terminal Value Estimation for Real-time and Multi-task Model Predictive Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Morita, Mitsuki, Yamamori, Satoshi, Yagi, Satoshi, Sugimoto, Norikazu, Morimoto, Jun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929533122248704
author Morita, Mitsuki
Yamamori, Satoshi
Yagi, Satoshi
Sugimoto, Norikazu
Morimoto, Jun
author_facet Morita, Mitsuki
Yamamori, Satoshi
Yagi, Satoshi
Sugimoto, Norikazu
Morimoto, Jun
contents While MPC enables nonlinear feedback control by solving an optimal control problem at each timestep, the computational burden tends to be significantly large, making it difficult to optimize a policy within the control period. To address this issue, one possible approach is to utilize terminal value learning to reduce computational costs. However, the learned value cannot be used for other tasks in situations where the task dynamically changes in the original MPC setup. In this study, we develop an MPC framework with goal-conditioned terminal value learning to achieve multitask policy optimization while reducing computational time. Furthermore, by using a hierarchical control structure that allows the upper-level trajectory planner to output appropriate goal-conditioned trajectories, we demonstrate that a robot model is able to generate diverse motions. We evaluate the proposed method on a bipedal inverted pendulum robot model and confirm that combining goal-conditioned terminal value learning with an upper-level trajectory planner enables real-time control; thus, the robot successfully tracks a target trajectory on sloped terrain.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04929
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Goal-Conditioned Terminal Value Estimation for Real-time and Multi-task Model Predictive Control
Morita, Mitsuki
Yamamori, Satoshi
Yagi, Satoshi
Sugimoto, Norikazu
Morimoto, Jun
Robotics
Machine Learning
Systems and Control
While MPC enables nonlinear feedback control by solving an optimal control problem at each timestep, the computational burden tends to be significantly large, making it difficult to optimize a policy within the control period. To address this issue, one possible approach is to utilize terminal value learning to reduce computational costs. However, the learned value cannot be used for other tasks in situations where the task dynamically changes in the original MPC setup. In this study, we develop an MPC framework with goal-conditioned terminal value learning to achieve multitask policy optimization while reducing computational time. Furthermore, by using a hierarchical control structure that allows the upper-level trajectory planner to output appropriate goal-conditioned trajectories, we demonstrate that a robot model is able to generate diverse motions. We evaluate the proposed method on a bipedal inverted pendulum robot model and confirm that combining goal-conditioned terminal value learning with an upper-level trajectory planner enables real-time control; thus, the robot successfully tracks a target trajectory on sloped terrain.
title Goal-Conditioned Terminal Value Estimation for Real-time and Multi-task Model Predictive Control
topic Robotics
Machine Learning
Systems and Control
url https://arxiv.org/abs/2410.04929