Fast Non-Episodic Adaptive Tuning of Robot Controllers with Online Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Preiss, James A., Xie, Fengze, Lin, Yiheng, Wierman, Adam, Yue, Yisong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911057394532352
author Preiss, James A.
Xie, Fengze
Lin, Yiheng
Wierman, Adam
Yue, Yisong
author_facet Preiss, James A.
Xie, Fengze
Lin, Yiheng
Wierman, Adam
Yue, Yisong
contents We study online algorithms to tune the parameters of a robot controller in a setting where the dynamics, policy class, and optimality objective are all time-varying. The system follows a single trajectory without episodes or state resets, and the time-varying information is not known in advance. Focusing on nonlinear geometric quadrotor controllers as a test case, we propose a practical implementation of a single-trajectory model-based online policy optimization algorithm, M-GAPS,along with reparameterizations of the quadrotor state space and policy class to improve the optimization landscape. In hardware experiments,we compare to model-based and model-free baselines that impose artificial episodes. We show that M-GAPS finds near-optimal parameters more quickly, especially when the episode length is not favorable. We also show that M-GAPS rapidly adapts to heavy unmodeled wind and payload disturbances, and achieves similar strong improvement on a 1:6-scale Ackermann-steered car. Our results demonstrate the hardware practicality of this emerging class of online policy optimization that offers significantly more flexibility than classic adaptive control, while being more stable and data-efficient than model-free reinforcement learning.
format Preprint
id arxiv_https___arxiv_org_abs_2507_10914
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fast Non-Episodic Adaptive Tuning of Robot Controllers with Online Policy Optimization
Preiss, James A.
Xie, Fengze
Lin, Yiheng
Wierman, Adam
Yue, Yisong
Robotics
Systems and Control
We study online algorithms to tune the parameters of a robot controller in a setting where the dynamics, policy class, and optimality objective are all time-varying. The system follows a single trajectory without episodes or state resets, and the time-varying information is not known in advance. Focusing on nonlinear geometric quadrotor controllers as a test case, we propose a practical implementation of a single-trajectory model-based online policy optimization algorithm, M-GAPS,along with reparameterizations of the quadrotor state space and policy class to improve the optimization landscape. In hardware experiments,we compare to model-based and model-free baselines that impose artificial episodes. We show that M-GAPS finds near-optimal parameters more quickly, especially when the episode length is not favorable. We also show that M-GAPS rapidly adapts to heavy unmodeled wind and payload disturbances, and achieves similar strong improvement on a 1:6-scale Ackermann-steered car. Our results demonstrate the hardware practicality of this emerging class of online policy optimization that offers significantly more flexibility than classic adaptive control, while being more stable and data-efficient than model-free reinforcement learning.
title Fast Non-Episodic Adaptive Tuning of Robot Controllers with Online Policy Optimization
topic Robotics
Systems and Control
url https://arxiv.org/abs/2507.10914