Quasi-Newton Compatible Actor-Critic for Deterministic Policies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kordabad, Arash Bahari, Brandner, Dean, Gros, Sebastien, Lucia, Sergio, Soudjani, Sadegh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908648515567616
author Kordabad, Arash Bahari
Brandner, Dean
Gros, Sebastien
Lucia, Sergio
Soudjani, Sadegh
author_facet Kordabad, Arash Bahari
Brandner, Dean
Gros, Sebastien
Lucia, Sergio
Soudjani, Sadegh
contents In this paper, we propose a second-order deterministic actor-critic framework in reinforcement learning that extends the classical deterministic policy gradient method to exploit curvature information of the performance function. Building on the concept of compatible function approximation for the critic, we introduce a quadratic critic that simultaneously preserves the true policy gradient and an approximation of the performance Hessian. A least-squares temporal difference learning scheme is then developed to estimate the quadratic critic parameters efficiently. This construction enables a quasi-Newton actor update using information learned by the critic, yielding faster convergence compared to first-order methods. The proposed approach is general and applicable to any differentiable policy class. Numerical examples demonstrate that the method achieves improved convergence and performance over standard deterministic actor-critic baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2511_09509
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Quasi-Newton Compatible Actor-Critic for Deterministic Policies
Kordabad, Arash Bahari
Brandner, Dean
Gros, Sebastien
Lucia, Sergio
Soudjani, Sadegh
Machine Learning
Systems and Control
Optimization and Control
In this paper, we propose a second-order deterministic actor-critic framework in reinforcement learning that extends the classical deterministic policy gradient method to exploit curvature information of the performance function. Building on the concept of compatible function approximation for the critic, we introduce a quadratic critic that simultaneously preserves the true policy gradient and an approximation of the performance Hessian. A least-squares temporal difference learning scheme is then developed to estimate the quadratic critic parameters efficiently. This construction enables a quasi-Newton actor update using information learned by the critic, yielding faster convergence compared to first-order methods. The proposed approach is general and applicable to any differentiable policy class. Numerical examples demonstrate that the method achieves improved convergence and performance over standard deterministic actor-critic baselines.
title Quasi-Newton Compatible Actor-Critic for Deterministic Policies
topic Machine Learning
Systems and Control
Optimization and Control
url https://arxiv.org/abs/2511.09509