DiffOP: Reinforcement Learning of Optimization-Based Control Policies via Implicit Policy Gradients

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bian, Yuexin, Feng, Jie, Shi, Yuanyuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915625931112448
author Bian, Yuexin
Feng, Jie
Shi, Yuanyuan
author_facet Bian, Yuexin
Feng, Jie
Shi, Yuanyuan
contents Real-world control systems require policies that are not only high-performing but also interpretable and robust. A promising direction toward this goal is model-based control, which learns system dynamics and cost functions from historical data and then uses these models to inform decision-making. Building on this paradigm, we introduce DiffOP, a novel framework for learning optimization-based control policies defined implicitly through optimization control problems. Without relying on value function approximation, DiffOP jointly learns the cost and dynamics models and directly optimizes the actual control costs using policy gradients. To enable this, we derive analytical policy gradients by applying implicit differentiation to the underlying optimization problem and integrating it with the standard policy gradient framework. Under standard regularity conditions, we establish that DiffOP converges to an $ε$-stationary point within $\mathcal{O}(ε^{-1})$ iterations. We demonstrate the effectiveness of DiffOP through experiments on nonlinear control tasks and power system voltage control with constraints. The code is available at https://github.com/alwaysbyx/DiffOP.
format Preprint
id arxiv_https___arxiv_org_abs_2411_07484
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DiffOP: Reinforcement Learning of Optimization-Based Control Policies via Implicit Policy Gradients
Bian, Yuexin
Feng, Jie
Shi, Yuanyuan
Systems and Control
Real-world control systems require policies that are not only high-performing but also interpretable and robust. A promising direction toward this goal is model-based control, which learns system dynamics and cost functions from historical data and then uses these models to inform decision-making. Building on this paradigm, we introduce DiffOP, a novel framework for learning optimization-based control policies defined implicitly through optimization control problems. Without relying on value function approximation, DiffOP jointly learns the cost and dynamics models and directly optimizes the actual control costs using policy gradients. To enable this, we derive analytical policy gradients by applying implicit differentiation to the underlying optimization problem and integrating it with the standard policy gradient framework. Under standard regularity conditions, we establish that DiffOP converges to an $ε$-stationary point within $\mathcal{O}(ε^{-1})$ iterations. We demonstrate the effectiveness of DiffOP through experiments on nonlinear control tasks and power system voltage control with constraints. The code is available at https://github.com/alwaysbyx/DiffOP.
title DiffOP: Reinforcement Learning of Optimization-Based Control Policies via Implicit Policy Gradients
topic Systems and Control
url https://arxiv.org/abs/2411.07484