Improving DAPO from a Mixed-Policy Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Tan, Hongze, Li, Yuchen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Block majorization-minimization with diminishing radius for constrained nonsmooth nonconvex optimization
by: Lyu, Hanbaek, et al.
Published: (2020)
by: Lyu, Hanbaek, et al.
Published: (2020)
Machine Learning Algorithms for Improving Exact Classical Solvers in Mixed Integer Continuous Optimization
by: Kimiaei, Morteza, et al.
Published: (2025)
by: Kimiaei, Morteza, et al.
Published: (2025)
Linear-Quadratic Mean-Field Reinforcement Learning: Convergence of Policy Gradient Methods
by: Carmona, René, et al.
Published: (2019)
by: Carmona, René, et al.
Published: (2019)
A Simple Mixture Policy Parameterization for Improving Sample Efficiency of CVaR Optimization
by: Luo, Yudong, et al.
Published: (2024)
by: Luo, Yudong, et al.
Published: (2024)
Empirical Risk Minimization with Shuffled SGD: A Primal-Dual Perspective and Improved Bounds
by: Cai, Xufeng, et al.
Published: (2023)
by: Cai, Xufeng, et al.
Published: (2023)
It's All in the Mix: Wasserstein Classification and Regression with Mixed Features
by: Belbasi, Reza, et al.
Published: (2023)
by: Belbasi, Reza, et al.
Published: (2023)
Wasserstein Formulation of Reinforcement Learning. An Optimal Transport Perspective on Policy Optimization
by: Dus, Mathias
Published: (2026)
by: Dus, Mathias
Published: (2026)
Revisiting LQR Control from the Perspective of Receding-Horizon Policy Gradient
by: Zhang, Xiangyuan, et al.
Published: (2023)
by: Zhang, Xiangyuan, et al.
Published: (2023)
Convergence and Complexity Guarantee for Inexact First-order Riemannian Optimization Algorithms
by: Li, Yuchen, et al.
Published: (2024)
by: Li, Yuchen, et al.
Published: (2024)
Convergence and complexity of block majorization-minimization for constrained block-Riemannian optimization
by: Li, Yuchen, et al.
Published: (2023)
by: Li, Yuchen, et al.
Published: (2023)
Elementary Analysis of Policy Gradient Methods
by: Liu, Jiacai, et al.
Published: (2024)
by: Liu, Jiacai, et al.
Published: (2024)
Learning to Price Bundles: A GCN Approach for Mixed Bundling
by: Ding, Liangyu, et al.
Published: (2025)
by: Ding, Liangyu, et al.
Published: (2025)
Active Learning Classification from a Signal Separation Perspective
by: Mhaskar, Hrushikesh, et al.
Published: (2025)
by: Mhaskar, Hrushikesh, et al.
Published: (2025)
On the Convergence of Policy in Unregularized Policy Mirror Descent
by: Lin, Dachao, et al.
Published: (2022)
by: Lin, Dachao, et al.
Published: (2022)
TRSVR: An Adaptive Stochastic Trust-Region Method with Variance Reduction
by: Fang, Yuchen, et al.
Published: (2026)
by: Fang, Yuchen, et al.
Published: (2026)
Learning to Stop: Deep Learning for Mean Field Optimal Stopping
by: Magnino, Lorenzo, et al.
Published: (2024)
by: Magnino, Lorenzo, et al.
Published: (2024)
On the Convergence of Policy Mirror Descent with Temporal Difference Evaluation
by: Liu, Jiacai, et al.
Published: (2025)
by: Liu, Jiacai, et al.
Published: (2025)
Scalable Mixed-Integer Optimization with Neural Constraints via Dual Decomposition
by: Zeng, Shuli, et al.
Published: (2025)
by: Zeng, Shuli, et al.
Published: (2025)
Policy Gradient Methods for Discrete Time Linear Quadratic Regulator With Random Parameters
by: Li, Deyue
Published: (2023)
by: Li, Deyue
Published: (2023)
Fair Generalized Linear Mixed Models
by: Burgard, Jan Pablo, et al.
Published: (2024)
by: Burgard, Jan Pablo, et al.
Published: (2024)
Fast Policy Learning for Linear Quadratic Control with Entropy Regularization
by: Guo, Xin, et al.
Published: (2023)
by: Guo, Xin, et al.
Published: (2023)
Improved Learning Rates for Stochastic Optimization
by: Li, Shaojie, et al.
Published: (2021)
by: Li, Shaojie, et al.
Published: (2021)
Policy Optimization in Hybrid Discrete-Continuous Action Spaces via Mixed Gradients
by: Alvo, Matias, et al.
Published: (2026)
by: Alvo, Matias, et al.
Published: (2026)
MiMuon: Mixed Muon Optimizer with Improved Generalization for Large Models
by: Huang, Feihu, et al.
Published: (2026)
by: Huang, Feihu, et al.
Published: (2026)
Provable Mixed-Noise Learning with Flow-Matching
by: Hagemann, Paul, et al.
Published: (2025)
by: Hagemann, Paul, et al.
Published: (2025)
Mixed-Integer Programming for Change-point Detection
by: Narula, Apoorva, et al.
Published: (2026)
by: Narula, Apoorva, et al.
Published: (2026)
Policy Gradient Converges to the Globally Optimal Policy for Nearly Linear-Quadratic Regulators
by: Han, Yinbin, et al.
Published: (2023)
by: Han, Yinbin, et al.
Published: (2023)
Constraint-Generation Policy Optimization (CGPO): Nonlinear Programming for Policy Optimization in Mixed Discrete-Continuous MDPs
by: Gimelfarb, Michael, et al.
Published: (2024)
by: Gimelfarb, Michael, et al.
Published: (2024)
Natural Policy Gradient as Doubly Smoothed Policy Iteration: A Bellman-Operator Framework
by: Nanda, Phalguni, et al.
Published: (2026)
by: Nanda, Phalguni, et al.
Published: (2026)
A Regularized Online Newton Method for Stochastic Convex Bandits with Linear Vanishing Noise
by: Zhan, Jingxin, et al.
Published: (2025)
by: Zhan, Jingxin, et al.
Published: (2025)
PolarGrad: A Class of Matrix-Gradient Optimizers from a Unifying Preconditioning Perspective
by: Lau, Tim Tsz-Kit, et al.
Published: (2025)
by: Lau, Tim Tsz-Kit, et al.
Published: (2025)
Mixed-feature Logistic Regression Robust to Distribution Shifts
by: Sun, Qingshi, et al.
Published: (2025)
by: Sun, Qingshi, et al.
Published: (2025)
Conformal Mixed-Integer Constraint Learning with Feasibility Guarantees
by: Ovalle, Daniel, et al.
Published: (2025)
by: Ovalle, Daniel, et al.
Published: (2025)
Responsible Machine Learning via Mixed-Integer Optimization
by: Justin, Nathan, et al.
Published: (2025)
by: Justin, Nathan, et al.
Published: (2025)
A Distance Metric for Mixed Integer Programming Instances
by: Maudet, Gwen, et al.
Published: (2025)
by: Maudet, Gwen, et al.
Published: (2025)
Hybrid Reinforcement Learning Framework for Mixed-Variable Problems
by: Zhai, Haoyan, et al.
Published: (2024)
by: Zhai, Haoyan, et al.
Published: (2024)
Tight Mixed-Integer Optimization Formulations for Prescriptive Trees
by: Biggs, Max, et al.
Published: (2023)
by: Biggs, Max, et al.
Published: (2023)
Soft Robust MDPs and Risk-Sensitive MDPs: Equivalence, Policy Gradient, and Sample Complexity
by: Zhang, Runyu, et al.
Published: (2023)
by: Zhang, Runyu, et al.
Published: (2023)
A Minibatch-SGD-Based Learning Meta-Policy for Inventory Systems with Myopic Optimal Policy
by: Lyu, Jiameng, et al.
Published: (2024)
by: Lyu, Jiameng, et al.
Published: (2024)
Double Duality: Variational Primal-Dual Policy Optimization for Constrained Reinforcement Learning
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
Similar Items
-
Block majorization-minimization with diminishing radius for constrained nonsmooth nonconvex optimization
by: Lyu, Hanbaek, et al.
Published: (2020) -
Machine Learning Algorithms for Improving Exact Classical Solvers in Mixed Integer Continuous Optimization
by: Kimiaei, Morteza, et al.
Published: (2025) -
Linear-Quadratic Mean-Field Reinforcement Learning: Convergence of Policy Gradient Methods
by: Carmona, René, et al.
Published: (2019) -
A Simple Mixture Policy Parameterization for Improving Sample Efficiency of CVaR Optimization
by: Luo, Yudong, et al.
Published: (2024) -
Empirical Risk Minimization with Shuffled SGD: A Primal-Dual Perspective and Improved Bounds
by: Cai, Xufeng, et al.
Published: (2023)