Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Zhongzhu, Yang, Yibo, Chen, Ziyan, Bie, Fengxiang, Xia, Haojun, Wu, Xiaoxia, Wu, Robert, Athiwaratkun, Ben, Ghanem, Bernard, Song, Shuaiwen Leon
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914018455715840
author Zhou, Zhongzhu
Yang, Yibo
Chen, Ziyan
Bie, Fengxiang
Xia, Haojun
Wu, Xiaoxia
Wu, Robert
Athiwaratkun, Ben
Ghanem, Bernard
Song, Shuaiwen Leon
author_facet Zhou, Zhongzhu
Yang, Yibo
Chen, Ziyan
Bie, Fengxiang
Xia, Haojun
Wu, Xiaoxia
Wu, Robert
Athiwaratkun, Ben
Ghanem, Bernard
Song, Shuaiwen Leon
contents Policy gradient (PG) methods in reinforcement learning frequently utilize deep neural networks (DNNs) to learn a shared backbone of feature representations used to compute likelihoods in an action selection layer. Numerous studies have been conducted on the convergence and global optima of policy networks, but few have analyzed representational structures of those underlying networks. While training an optimal policy DNN, we observed that under certain constraints, a gentle structure resembling neural collapse, which we refer to as Action Collapse (AC), emerges. This suggests that 1) the state-action activations (i.e. last-layer features) sharing the same optimal actions collapse towards those optimal actions respective mean activations; 2) the variability of activations sharing the same optimal actions converges to zero; 3) the weights of action selection layer and the mean activations collapse to a simplex equiangular tight frame (ETF). Our early work showed those aforementioned constraints to be necessary for these observations. Since the collapsed ETF of optimal policy DNNs maximally separates the pair-wise angles of all actions in the state-action space, we naturally raise a question: can we learn an optimal policy using an ETF structure as a (fixed) target configuration in the action selection layer? Our analytical proof shows that learning activations with a fixed ETF as action selection layer naturally leads to the AC. We thus propose the Action Collapse Policy Gradient (ACPG) method, which accordingly affixes a synthetic ETF as our action selection layer. ACPG induces the policy DNN to produce such an ideal configuration in the action selection layer while remaining optimal. Our experiments across various OpenAI Gym environments demonstrate that our technique can be integrated into any discrete PG methods and lead to favorable reward improvements more quickly and robustly.
format Preprint
id arxiv_https___arxiv_org_abs_2509_02737
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient
Zhou, Zhongzhu
Yang, Yibo
Chen, Ziyan
Bie, Fengxiang
Xia, Haojun
Wu, Xiaoxia
Wu, Robert
Athiwaratkun, Ben
Ghanem, Bernard
Song, Shuaiwen Leon
Machine Learning
Policy gradient (PG) methods in reinforcement learning frequently utilize deep neural networks (DNNs) to learn a shared backbone of feature representations used to compute likelihoods in an action selection layer. Numerous studies have been conducted on the convergence and global optima of policy networks, but few have analyzed representational structures of those underlying networks. While training an optimal policy DNN, we observed that under certain constraints, a gentle structure resembling neural collapse, which we refer to as Action Collapse (AC), emerges. This suggests that 1) the state-action activations (i.e. last-layer features) sharing the same optimal actions collapse towards those optimal actions respective mean activations; 2) the variability of activations sharing the same optimal actions converges to zero; 3) the weights of action selection layer and the mean activations collapse to a simplex equiangular tight frame (ETF). Our early work showed those aforementioned constraints to be necessary for these observations. Since the collapsed ETF of optimal policy DNNs maximally separates the pair-wise angles of all actions in the state-action space, we naturally raise a question: can we learn an optimal policy using an ETF structure as a (fixed) target configuration in the action selection layer? Our analytical proof shows that learning activations with a fixed ETF as action selection layer naturally leads to the AC. We thus propose the Action Collapse Policy Gradient (ACPG) method, which accordingly affixes a synthetic ETF as our action selection layer. ACPG induces the policy DNN to produce such an ideal configuration in the action selection layer while remaining optimal. Our experiments across various OpenAI Gym environments demonstrate that our technique can be integrated into any discrete PG methods and lead to favorable reward improvements more quickly and robustly.
title Imitate Optimal Policy: Prevail and Induce Action Collapse in Policy Gradient
topic Machine Learning
url https://arxiv.org/abs/2509.02737