Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Sun, Mingfei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911694773551104
author Sun, Mingfei
author_facet Sun, Mingfei
contents Natural policy gradients improve optimization by accounting for the geometry of distribution space, but their practical use is limited by the cost of estimating and inverting the Fisher matrix. We present Randomized Advantage Transformation (RAT), a method for estimating Tikhonov-regularized natural policy gradients via direct backpropagation. By applying the Woodbury formula, we reformulate the regularized natural policy gradients as vanilla policy gradients with a transformed advantage. RAT computes this transformation efficiently via randomized block Kaczmarz iterations on on-policy mini-batches, avoiding explicit Fisher construction, conjugate-gradient solvers, and architecture-specific approximations. We provide convergence guarantees for RAT and demonstrate empirically that it matches or exceeds established natural-gradient methods across continuous and visual control benchmarks, while remaining simple to implement and compatible with various architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18591
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation
Sun, Mingfei
Machine Learning
Artificial Intelligence
Natural policy gradients improve optimization by accounting for the geometry of distribution space, but their practical use is limited by the cost of estimating and inverting the Fisher matrix. We present Randomized Advantage Transformation (RAT), a method for estimating Tikhonov-regularized natural policy gradients via direct backpropagation. By applying the Woodbury formula, we reformulate the regularized natural policy gradients as vanilla policy gradients with a transformed advantage. RAT computes this transformation efficiently via randomized block Kaczmarz iterations on on-policy mini-batches, avoiding explicit Fisher construction, conjugate-gradient solvers, and architecture-specific approximations. We provide convergence guarantees for RAT and demonstrate empirically that it matches or exceeds established natural-gradient methods across continuous and visual control benchmarks, while remaining simple to implement and compatible with various architectures.
title Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.18591