GPG: Generalized Policy Gradient Theorem for Transformer-based Policies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mao, Hangyu, Dong, Guangting, Dou, Zhicheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915668469743616
author Mao, Hangyu
Dong, Guangting
Dou, Zhicheng
author_facet Mao, Hangyu
Dong, Guangting
Dou, Zhicheng
contents We present the Generalized Policy Gradient (GPG) Theorem, specifically designed for Transformer-based policies. Notably, we demonstrate that both standard Policy Gradient Theorem and GRPO emerge as special cases within our GPG framework. Furthermore, we explore its practical applications in training Large Language Models (LLMs), offering new insights into efficient policy optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10365
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GPG: Generalized Policy Gradient Theorem for Transformer-based Policies
Mao, Hangyu
Dong, Guangting
Dou, Zhicheng
Machine Learning
Artificial Intelligence
Computation and Language
We present the Generalized Policy Gradient (GPG) Theorem, specifically designed for Transformer-based policies. Notably, we demonstrate that both standard Policy Gradient Theorem and GRPO emerge as special cases within our GPG framework. Furthermore, we explore its practical applications in training Large Language Models (LLMs), offering new insights into efficient policy optimization.
title GPG: Generalized Policy Gradient Theorem for Transformer-based Policies
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2512.10365