Bi-Level Policy Optimization with Nyström Hypergradients

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Prakash, Arjun, He, Naicheng, Goktas, Denizalp, Makar-Limanov, Jacob, Greenwald, Amy
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912971538563072
author Prakash, Arjun
He, Naicheng
Goktas, Denizalp
Makar-Limanov, Jacob
Greenwald, Amy
author_facet Prakash, Arjun
He, Naicheng
Goktas, Denizalp
Makar-Limanov, Jacob
Greenwald, Amy
contents The dependency of the actor on the critic in actor-critic (AC) reinforcement learning means that AC can be characterized as a bilevel optimization (BLO) problem, also called a Stackelberg game. This characterization motivates two modifications to vanilla AC algorithms. First, the critic's update should be nested to learn a best response to the actor's policy. Second, the actor should update according to a hypergradient that takes changes in the critic's behavior into account. Computing this hypergradient involves finding an inverse Hessian vector product, a process that can be numerically unstable. We thus propose a new algorithm, Bilevel Policy Optimization with Nyström Hypergradients (BLPO), which uses nesting to account for the nested structure of BLO, and leverages the Nyström method to compute the hypergradient. Theoretically, we prove BLPO converges to (a point that satisfies the necessary conditions for) a local strong Stackelberg equilibrium in polynomial time with high probability, assuming a linear parametrization of the critic's objective. Empirically, we demonstrate that BLPO performs on par with or better than PPO on a variety of discrete and continuous control tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11714
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bi-Level Policy Optimization with Nyström Hypergradients
Prakash, Arjun
He, Naicheng
Goktas, Denizalp
Makar-Limanov, Jacob
Greenwald, Amy
Machine Learning
Artificial Intelligence
Computer Science and Game Theory
The dependency of the actor on the critic in actor-critic (AC) reinforcement learning means that AC can be characterized as a bilevel optimization (BLO) problem, also called a Stackelberg game. This characterization motivates two modifications to vanilla AC algorithms. First, the critic's update should be nested to learn a best response to the actor's policy. Second, the actor should update according to a hypergradient that takes changes in the critic's behavior into account. Computing this hypergradient involves finding an inverse Hessian vector product, a process that can be numerically unstable. We thus propose a new algorithm, Bilevel Policy Optimization with Nyström Hypergradients (BLPO), which uses nesting to account for the nested structure of BLO, and leverages the Nyström method to compute the hypergradient. Theoretically, we prove BLPO converges to (a point that satisfies the necessary conditions for) a local strong Stackelberg equilibrium in polynomial time with high probability, assuming a linear parametrization of the critic's objective. Empirically, we demonstrate that BLPO performs on par with or better than PPO on a variety of discrete and continuous control tasks.
title Bi-Level Policy Optimization with Nyström Hypergradients
topic Machine Learning
Artificial Intelligence
Computer Science and Game Theory
url https://arxiv.org/abs/2505.11714