GPG: Generalized Policy Gradient Theorem for Transformer-based Policies
Fuente:
arXiv
Saved in:
| Main Authors: | Mao, Hangyu, Dong, Guangting, Dou, Zhicheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agentic Reinforced Policy Optimization
by: Dong, Guanting, et al.
Published: (2025)
by: Dong, Guanting, et al.
Published: (2025)
Agentic Entropy-Balanced Policy Optimization
by: Dong, Guanting, et al.
Published: (2025)
by: Dong, Guanting, et al.
Published: (2025)
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
by: Dong, Guanting, et al.
Published: (2025)
by: Dong, Guanting, et al.
Published: (2025)
Simple Policy Gradients for Reasoning with Diffusion Language Models
by: Zhan, Anthony
Published: (2025)
by: Zhan, Anthony
Published: (2025)
Towards Mixed-Modal Retrieval for Universal Retrieval-Augmented Generation
by: Zhang, Chenghao, et al.
Published: (2025)
by: Zhang, Chenghao, et al.
Published: (2025)
DCPO: Dynamic Clipping Policy Optimization
by: Yang, Shihui, et al.
Published: (2025)
by: Yang, Shihui, et al.
Published: (2025)
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
by: Su, Zhenpeng, et al.
Published: (2025)
by: Su, Zhenpeng, et al.
Published: (2025)
Fine-Tuning Discrete Diffusion Models with Policy Gradient Methods
by: Zekri, Oussama, et al.
Published: (2025)
by: Zekri, Oussama, et al.
Published: (2025)
On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Interpreting and Controlling LLM Reasoning through Integrated Policy Gradient
by: Li, Changming, et al.
Published: (2026)
by: Li, Changming, et al.
Published: (2026)
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation
by: Dong, Guanting, et al.
Published: (2024)
by: Dong, Guanting, et al.
Published: (2024)
GTPO: Stabilizing Group Relative Policy Optimization via Gradient and Entropy Control
by: Simoni, Marco, et al.
Published: (2025)
by: Simoni, Marco, et al.
Published: (2025)
Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models
by: Li, Chengao, et al.
Published: (2025)
by: Li, Chengao, et al.
Published: (2025)
Policy-Gradient Training of Language Models for Ranking
by: Gao, Ge, et al.
Published: (2023)
by: Gao, Ge, et al.
Published: (2023)
Toward General Instruction-Following Alignment for Retrieval-Augmented Generation
by: Dong, Guanting, et al.
Published: (2024)
by: Dong, Guanting, et al.
Published: (2024)
Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space
by: Li, Hengli, et al.
Published: (2025)
by: Li, Hengli, et al.
Published: (2025)
EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis
by: Song, Xiaoshuai, et al.
Published: (2026)
by: Song, Xiaoshuai, et al.
Published: (2026)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
by: Cui, Sijia, et al.
Published: (2026)
by: Cui, Sijia, et al.
Published: (2026)
C$^2$GSPG: Confidence-calibrated Group Sequence Policy Gradient towards Self-aware Reasoning
by: Liu, Haotian, et al.
Published: (2025)
by: Liu, Haotian, et al.
Published: (2025)
Agentic Policy Optimization via Instruction-Policy Co-Evolution
by: Zhou, Han, et al.
Published: (2025)
by: Zhou, Han, et al.
Published: (2025)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
by: Yang, Wenkai, et al.
Published: (2026)
by: Yang, Wenkai, et al.
Published: (2026)
Fibration Policy Optimization
by: Li, Chang, et al.
Published: (2026)
by: Li, Chang, et al.
Published: (2026)
Adaptive Social Learning via Mode Policy Optimization for Language Agents
by: Wang, Minzheng, et al.
Published: (2025)
by: Wang, Minzheng, et al.
Published: (2025)
ReDit: Reward Dithering for Improved LLM Policy Optimization
by: Wei, Chenxing, et al.
Published: (2025)
by: Wei, Chenxing, et al.
Published: (2025)
Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks
by: Guo, Ruohao, et al.
Published: (2025)
by: Guo, Ruohao, et al.
Published: (2025)
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
by: Agarwal, Rishabh, et al.
Published: (2023)
by: Agarwal, Rishabh, et al.
Published: (2023)
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models
by: Xu, Xingcheng
Published: (2025)
by: Xu, Xingcheng
Published: (2025)
Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning
by: Liu, Rui, et al.
Published: (2025)
by: Liu, Rui, et al.
Published: (2025)
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training
by: Pan, Chengjun, et al.
Published: (2026)
by: Pan, Chengjun, et al.
Published: (2026)
Soft Adaptive Policy Optimization
by: Gao, Chang, et al.
Published: (2025)
by: Gao, Chang, et al.
Published: (2025)
Group Sequence Policy Optimization
by: Zheng, Chujie, et al.
Published: (2025)
by: Zheng, Chujie, et al.
Published: (2025)
Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning
by: Zhao, Shiwan, et al.
Published: (2026)
by: Zhao, Shiwan, et al.
Published: (2026)
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier
by: Jiang, Yuhua, et al.
Published: (2025)
by: Jiang, Yuhua, et al.
Published: (2025)
Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning
by: Wang, Kaiwen, et al.
Published: (2024)
by: Wang, Kaiwen, et al.
Published: (2024)
The Condensate Theorem: Transformers are O(n), Not $O(n^2)$
by: Williams, Jorge L. Ruiz
Published: (2026)
by: Williams, Jorge L. Ruiz
Published: (2026)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
Policies and Evaluation for Online Meeting Summarization
by: Schneider, Felix, et al.
Published: (2025)
by: Schneider, Felix, et al.
Published: (2025)
Causally-Enhanced Reinforcement Policy Optimization
by: Wang, Xiangqi, et al.
Published: (2025)
by: Wang, Xiangqi, et al.
Published: (2025)
COPO: Consistency-Aware Policy Optimization
by: Han, Jinghang, et al.
Published: (2025)
by: Han, Jinghang, et al.
Published: (2025)
Policy Learning with a Language Bottleneck
by: Srivastava, Megha, et al.
Published: (2024)
by: Srivastava, Megha, et al.
Published: (2024)
Similar Items
-
Agentic Reinforced Policy Optimization
by: Dong, Guanting, et al.
Published: (2025) -
Agentic Entropy-Balanced Policy Optimization
by: Dong, Guanting, et al.
Published: (2025) -
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
by: Dong, Guanting, et al.
Published: (2025) -
Simple Policy Gradients for Reasoning with Diffusion Language Models
by: Zhan, Anthony
Published: (2025) -
Towards Mixed-Modal Retrieval for Universal Retrieval-Augmented Generation
by: Zhang, Chenghao, et al.
Published: (2025)