Policy Learning for Perturbance-wise Linear Quadratic Control Problem

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Haoran, Zhang, Wenhao, Wu, Xianping
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915610173112320
author Zhang, Haoran
Zhang, Wenhao
Wu, Xianping
author_facet Zhang, Haoran
Zhang, Wenhao
Wu, Xianping
contents We study finite horizon linear quadratic control with additive noise in a perturbancewise framework that unifies the classical model, a constraint embedded affine policy class, and a distributionally robust formulation with a Wasserstein ambiguity set. Based on an augmented affine representation, we model feasibility as an affine perturbation and unknown noise as distributional perturbation from samples, thereby addressing constrained implementation and model uncertainty in a single scheme. First, we construct an implementable policy gradient method that accommodates nonzero noise means estimated from data. Second, we analyze its convergence under constant stepsizes chosen as simple polynomials of problem parameters, ensuring global decrease of the value function. Finally, numerical studies: mean variance portfolio allocation and dynamic benchmark tracking on real data, validating stable convergence and illuminating sensitivity tradeoffs across horizon length, trading cost intensity, state penalty scale, and estimation window.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07388
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Policy Learning for Perturbance-wise Linear Quadratic Control Problem
Zhang, Haoran
Zhang, Wenhao
Wu, Xianping
Optimization and Control
We study finite horizon linear quadratic control with additive noise in a perturbancewise framework that unifies the classical model, a constraint embedded affine policy class, and a distributionally robust formulation with a Wasserstein ambiguity set. Based on an augmented affine representation, we model feasibility as an affine perturbation and unknown noise as distributional perturbation from samples, thereby addressing constrained implementation and model uncertainty in a single scheme. First, we construct an implementable policy gradient method that accommodates nonzero noise means estimated from data. Second, we analyze its convergence under constant stepsizes chosen as simple polynomials of problem parameters, ensuring global decrease of the value function. Finally, numerical studies: mean variance portfolio allocation and dynamic benchmark tracking on real data, validating stable convergence and illuminating sensitivity tradeoffs across horizon length, trading cost intensity, state penalty scale, and estimation window.
title Policy Learning for Perturbance-wise Linear Quadratic Control Problem
topic Optimization and Control
url https://arxiv.org/abs/2511.07388