GDEPO: Group Dual-dynamic and Equal-right Advantage Policy Optimization with Enhanced Training Data Utilization for Sample-Constrained Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Zhengqing, Liu, Xinyang, Zhang, Yi, Guo, Fan, Jia, ChengXun, Wan, Junchen, Liu, Yao, Liu, Qi, Huang, Jihao, Song, Kang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GAGPO: Generalized Advantage Grouped Policy Optimization
by: Zhu, Siyuan, et al.
Published: (2026)
by: Zhu, Siyuan, et al.
Published: (2026)
Dual complexes of qdlt Fano type models and strong complete regularity
by: Liu, Jihao, et al.
Published: (2026)
by: Liu, Jihao, et al.
Published: (2026)
Double Duality: Variational Primal-Dual Policy Optimization for Constrained Reinforcement Learning
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
Policy Split: Incentivizing Dual-Mode Exploration in LLM Reinforcement with Dual-Mode Entropy Regularization
by: Yao, Jiashu, et al.
Published: (2026)
by: Yao, Jiashu, et al.
Published: (2026)
Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement Learning
by: Liu, Tenglong, et al.
Published: (2024)
by: Liu, Tenglong, et al.
Published: (2024)
Let the Results Speak: A Replication-First Paradigm for LLM Behavioral Benchmarking
by: Yuming, et al.
Published: (2026)
by: Yuming, et al.
Published: (2026)
On a question of Kollár and Kovács
by: Liu, Jihao
Published: (2026)
by: Liu, Jihao
Published: (2026)
On a question of Mauri and Moraga
by: Liu, Jihao
Published: (2026)
by: Liu, Jihao
Published: (2026)
A question on klt type varieties of Han and Jiang
by: Liu, Jihao
Published: (2026)
by: Liu, Jihao
Published: (2026)
An example of a very non-movable effective divisor
by: Liu, Jihao
Published: (2026)
by: Liu, Jihao
Published: (2026)
Group-in-Group Policy Optimization for LLM Agent Training
by: Feng, Lang, et al.
Published: (2025)
by: Feng, Lang, et al.
Published: (2025)
Sample-Efficient Policy Constraint Offline Deep Reinforcement Learning based on Sample Filtering
by: Chen, Yuanhao, et al.
Published: (2025)
by: Chen, Yuanhao, et al.
Published: (2025)
Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction
by: Ding, Fei, et al.
Published: (2026)
by: Ding, Fei, et al.
Published: (2026)
The minimal volume of surfaces of log general type with non-empty non-klt locus
by: Liu, Jihao, et al.
Published: (2023)
by: Liu, Jihao, et al.
Published: (2023)
The minimal volume of stable surfaces of rank one
by: Liu, Jihao, et al.
Published: (2026)
by: Liu, Jihao, et al.
Published: (2026)
ACC for local volumes
by: Han, Jingjun, et al.
Published: (2024)
by: Han, Jingjun, et al.
Published: (2024)
Outcome-Aware Tool Selection for Semantic Routers: Latency-Constrained Learning Without LLM Inference
by: Chen, Huamin, et al.
Published: (2026)
by: Chen, Huamin, et al.
Published: (2026)
On the Cause of Unfairness: A Training Sample Perspective
by: Yao, Yuanshun, et al.
Published: (2023)
by: Yao, Yuanshun, et al.
Published: (2023)
MAPO: Mixed Advantage Policy Optimization
by: Huang, Wenke, et al.
Published: (2025)
by: Huang, Wenke, et al.
Published: (2025)
Unified Entropy Optimization for Open-Set Test-Time Adaptation
by: Gao, Zhengqing, et al.
Published: (2024)
by: Gao, Zhengqing, et al.
Published: (2024)
Direct Advantage Regression: Aligning LLMs with Online AI Reward
by: He, Li, et al.
Published: (2025)
by: He, Li, et al.
Published: (2025)
Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation
by: He, Xixiang, et al.
Published: (2026)
by: He, Xixiang, et al.
Published: (2026)
Classification of threefold enc cDV quotient singularities
by: Han, Jingjun, et al.
Published: (2025)
by: Han, Jingjun, et al.
Published: (2025)
Non-algebraicity of non-abundant foliations and abundance for adjoint foliated structures
by: Liu, Jihao, et al.
Published: (2025)
by: Liu, Jihao, et al.
Published: (2025)
On termination of flips and exceptionally non-canonical singularities
by: Han, Jingjun, et al.
Published: (2022)
by: Han, Jingjun, et al.
Published: (2022)
A generalized non-vanishing theorem on surfaces
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
Existence of the minimal model program for log canonical generalized pairs
by: Hu, Zhengyu, et al.
Published: (2026)
by: Hu, Zhengyu, et al.
Published: (2026)
Non-vanishing implies numerical dimension one abundance
by: Liu, Jihao, et al.
Published: (2025)
by: Liu, Jihao, et al.
Published: (2025)
Shokurov's global index conjecture for threefold foliations
by: Liu, Jihao, et al.
Published: (2026)
by: Liu, Jihao, et al.
Published: (2026)
Space Group Constrained Crystal Generation
by: Jiao, Rui, et al.
Published: (2024)
by: Jiao, Rui, et al.
Published: (2024)
Adversarial Samples Are Not Created Equal
by: Crawford, Jennifer, et al.
Published: (2026)
by: Crawford, Jennifer, et al.
Published: (2026)
When is Offline Policy Selection Sample Efficient for Reinforcement Learning?
by: Liu, Vincent, et al.
Published: (2023)
by: Liu, Vincent, et al.
Published: (2023)
Long time well-posedness for the 3D Prandtl boundary layer equations with a special structure
by: Qin, Yuming, et al.
Published: (2024)
by: Qin, Yuming, et al.
Published: (2024)
Local-in-time well-posedness for 2D compressible magneto-micropolar boundary layer in Sobolev spaces
by: Qin, Yuming, et al.
Published: (2025)
by: Qin, Yuming, et al.
Published: (2025)
Beyond One-Way Pruning: Bidirectional Pruning-Regrowth for Extreme Accuracy-Sparsity Tradeoff
by: Liu, Junchen, et al.
Published: (2025)
by: Liu, Junchen, et al.
Published: (2025)
Adaptive Group Policy Optimization: Towards Stable Training and Token-Efficient Reasoning
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
Constrained Group Relative Policy Optimization
by: Girgis, Roger, et al.
Published: (2026)
by: Girgis, Roger, et al.
Published: (2026)
Test-Time Training with KV Binding Is Secretly Linear Attention
by: Liu, Junchen, et al.
Published: (2026)
by: Liu, Junchen, et al.
Published: (2026)
Empirical Study of Named Entity Recognition Performance Using Distribution-aware Word Embedding
by: Chen, Xin, et al.
Published: (2021)
by: Chen, Xin, et al.
Published: (2021)
Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage
by: Jin, Weiqiang, et al.
Published: (2025)
by: Jin, Weiqiang, et al.
Published: (2025)
Similar Items
-
GAGPO: Generalized Advantage Grouped Policy Optimization
by: Zhu, Siyuan, et al.
Published: (2026) -
Dual complexes of qdlt Fano type models and strong complete regularity
by: Liu, Jihao, et al.
Published: (2026) -
Double Duality: Variational Primal-Dual Policy Optimization for Constrained Reinforcement Learning
by: Li, Zihao, et al.
Published: (2024) -
Policy Split: Incentivizing Dual-Mode Exploration in LLM Reinforcement with Dual-Mode Entropy Regularization
by: Yao, Jiashu, et al.
Published: (2026) -
Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement Learning
by: Liu, Tenglong, et al.
Published: (2024)