Divergence-Augmented Policy Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Qing, Li, Yingru, Xiong, Jiechao, Zhang, Tong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Logit Dynamics in Softmax Policy Gradient Methods
von: Li, Yingru
Veröffentlicht: (2025)
von: Li, Yingru
Veröffentlicht: (2025)
Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models
von: Fan, Jiajun, et al.
Veröffentlicht: (2025)
von: Fan, Jiajun, et al.
Veröffentlicht: (2025)
Beyond KL Divergence: Policy Optimization with Flexible Bregman Divergences for LLM Reasoning
von: Yuan, Rui, et al.
Veröffentlicht: (2026)
von: Yuan, Rui, et al.
Veröffentlicht: (2026)
Beyond Precision: Training-Inference Mismatch is an Optimization Problem and Simple LR Scheduling Fixes It
von: Zhang, Yaxiang, et al.
Veröffentlicht: (2026)
von: Zhang, Yaxiang, et al.
Veröffentlicht: (2026)
StaRPO: Stability-Augmented Reinforcement Policy Optimization
von: Zhang, Jinghan, et al.
Veröffentlicht: (2026)
von: Zhang, Jinghan, et al.
Veröffentlicht: (2026)
APO: Alpha-Divergence Preference Optimization
von: Zixian, Wang
Veröffentlicht: (2025)
von: Zixian, Wang
Veröffentlicht: (2025)
Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning
von: Yang, Hanlin, et al.
Veröffentlicht: (2024)
von: Yang, Hanlin, et al.
Veröffentlicht: (2024)
A Note on Hybrid Online Reinforcement and Imitation Learning for LLMs: Formulations and Algorithms
von: Li, Yingru, et al.
Veröffentlicht: (2025)
von: Li, Yingru, et al.
Veröffentlicht: (2025)
Dynamic Vocabulary Pruning: Stable LLM-RL by Taming the Tail
von: Li, Yingru, et al.
Veröffentlicht: (2025)
von: Li, Yingru, et al.
Veröffentlicht: (2025)
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
von: Liu, Zeyuan, et al.
Veröffentlicht: (2026)
von: Liu, Zeyuan, et al.
Veröffentlicht: (2026)
Q-Star Meets Scalable Posterior Sampling: Bridging Theory and Practice via HyperAgent
von: Li, Yingru, et al.
Veröffentlicht: (2024)
von: Li, Yingru, et al.
Veröffentlicht: (2024)
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
von: Zhao, Qingyue, et al.
Veröffentlicht: (2025)
von: Zhao, Qingyue, et al.
Veröffentlicht: (2025)
A Unified Framework for Rethinking Policy Divergence Measures in GRPO
von: Wu, Qingyuan, et al.
Veröffentlicht: (2026)
von: Wu, Qingyuan, et al.
Veröffentlicht: (2026)
UCPO: Uncertainty-Aware Policy Optimization
von: Zeng, Xianzhou, et al.
Veröffentlicht: (2026)
von: Zeng, Xianzhou, et al.
Veröffentlicht: (2026)
Trust Region Masking for Long-Horizon LLM Reinforcement Learning
von: Li, Yingru, et al.
Veröffentlicht: (2025)
von: Li, Yingru, et al.
Veröffentlicht: (2025)
ESPO: Early-Stopping Proximal Policy Optimization
von: Li, Zihang, et al.
Veröffentlicht: (2026)
von: Li, Zihang, et al.
Veröffentlicht: (2026)
Unearthing Gems from Stones: Policy Optimization with Negative Sample Augmentation for LLM Reasoning
von: Yang, Zhaohui, et al.
Veröffentlicht: (2025)
von: Yang, Zhaohui, et al.
Veröffentlicht: (2025)
Prior-dependent analysis of posterior sampling reinforcement learning with function approximation
von: Li, Yingru, et al.
Veröffentlicht: (2024)
von: Li, Yingru, et al.
Veröffentlicht: (2024)
Expert Divergence Learning for MoE-based Language Models
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
Orthogonalized Policy Optimization:Policy Optimization as Orthogonal Projection in Hilbert Space
von: Zixian, Wang
Veröffentlicht: (2026)
von: Zixian, Wang
Veröffentlicht: (2026)
Reparameterization Proximal Policy Optimization
von: Zhong, Hai, et al.
Veröffentlicht: (2025)
von: Zhong, Hai, et al.
Veröffentlicht: (2025)
NGRPO: Negative-enhanced Group Relative Policy Optimization
von: Nan, Gongrui, et al.
Veröffentlicht: (2025)
von: Nan, Gongrui, et al.
Veröffentlicht: (2025)
Reparameterization Flow Policy Optimization
von: Zhong, Hai, et al.
Veröffentlicht: (2026)
von: Zhong, Hai, et al.
Veröffentlicht: (2026)
Relative Policy-Transition Optimization for Fast Policy Transfer
von: Xu, Jiawei, et al.
Veröffentlicht: (2022)
von: Xu, Jiawei, et al.
Veröffentlicht: (2022)
SPOT: Scalable Policy Optimization with Trees for Markov Decision Processes
von: Xiong, Xuyuan, et al.
Veröffentlicht: (2025)
von: Xiong, Xuyuan, et al.
Veröffentlicht: (2025)
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
Rényi Divergence Deep Mutual Learning
von: Huang, Weipeng, et al.
Veröffentlicht: (2022)
von: Huang, Weipeng, et al.
Veröffentlicht: (2022)
Federated Neural Architecture Search with Model-Agnostic Meta Learning
von: Huang, Xinyuan, et al.
Veröffentlicht: (2025)
von: Huang, Xinyuan, et al.
Veröffentlicht: (2025)
Group Orthogonalized Policy Optimization:Group Policy Optimization as Orthogonal Projection in Hilbert Space
von: Zixian, Wang
Veröffentlicht: (2026)
von: Zixian, Wang
Veröffentlicht: (2026)
From Atoms to Chains: Divergence-Guided Reasoning Curriculum for Unlabeled LLM Domain Adaptation
von: Wang, Yongqi, et al.
Veröffentlicht: (2026)
von: Wang, Yongqi, et al.
Veröffentlicht: (2026)
GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2025)
Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective
von: Zhang, Yuheng, et al.
Veröffentlicht: (2026)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2026)
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
von: Huang, Chenghua, et al.
Veröffentlicht: (2025)
von: Huang, Chenghua, et al.
Veröffentlicht: (2025)
On-Policy Optimization of ANFIS Policies Using Proximal Policy Optimization
von: Shankar, Kaaustaaub, et al.
Veröffentlicht: (2025)
von: Shankar, Kaaustaaub, et al.
Veröffentlicht: (2025)
Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
von: Li, Xuan, et al.
Veröffentlicht: (2026)
von: Li, Xuan, et al.
Veröffentlicht: (2026)
The Optimal Token Baseline: Variance Reduction for Long-Horizon LLM-RL
von: Li, Yingru, et al.
Veröffentlicht: (2026)
von: Li, Yingru, et al.
Veröffentlicht: (2026)
Optimal Stability of KL Divergence under Gaussian Perturbations
von: Pan, Jialu, et al.
Veröffentlicht: (2026)
von: Pan, Jialu, et al.
Veröffentlicht: (2026)
Fairness in Graph Learning Augmented with Machine Learning: A Survey
von: Luo, Renqiang, et al.
Veröffentlicht: (2025)
von: Luo, Renqiang, et al.
Veröffentlicht: (2025)
Calibration-Aware Policy Optimization for Reasoning LLMs
von: Wang, Ziqi, et al.
Veröffentlicht: (2026)
von: Wang, Ziqi, et al.
Veröffentlicht: (2026)
Fractal Landscapes in Policy Optimization
von: Wang, Tao, et al.
Veröffentlicht: (2023)
von: Wang, Tao, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Logit Dynamics in Softmax Policy Gradient Methods
von: Li, Yingru
Veröffentlicht: (2025) -
Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models
von: Fan, Jiajun, et al.
Veröffentlicht: (2025) -
Beyond KL Divergence: Policy Optimization with Flexible Bregman Divergences for LLM Reasoning
von: Yuan, Rui, et al.
Veröffentlicht: (2026) -
Beyond Precision: Training-Inference Mismatch is an Optimization Problem and Simple LR Scheduling Fixes It
von: Zhang, Yaxiang, et al.
Veröffentlicht: (2026) -
StaRPO: Stability-Augmented Reinforcement Policy Optimization
von: Zhang, Jinghan, et al.
Veröffentlicht: (2026)