Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Jiashun, Obando-Ceron, Johan, Lu, Han, He, Yancheng, Wang, Weixun, Su, Wenbo, Zheng, Bo, Castro, Pablo Samuel, Courville, Aaron, Pan, Ling |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning
by: Liu, Jiashun, et al.
Published: (2025)
by: Liu, Jiashun, et al.
Published: (2025)
Neuroplastic Expansion in Deep Reinforcement Learning
by: Liu, Jiashun, et al.
Published: (2024)
by: Liu, Jiashun, et al.
Published: (2024)
Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learning
by: Liu, Jiashun, et al.
Published: (2025)
by: Liu, Jiashun, et al.
Published: (2025)
In value-based deep reinforcement learning, a pruned network is a good network
by: Obando-Ceron, Johan, et al.
Published: (2024)
by: Obando-Ceron, Johan, et al.
Published: (2024)
The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning Networks
by: Mayor, Walter, et al.
Published: (2025)
by: Mayor, Walter, et al.
Published: (2025)
On the consistency of hyper-parameter selection in value-based deep reinforcement learning
by: Obando-Ceron, Johan, et al.
Published: (2024)
by: Obando-Ceron, Johan, et al.
Published: (2024)
Don't flatten, tokenize! Unlocking the key to SoftMoE's efficacy in deep RL
by: Sokar, Ghada, et al.
Published: (2024)
by: Sokar, Ghada, et al.
Published: (2024)
Mitigating Plasticity Loss in Continual Reinforcement Learning by Reducing Churn
by: Tang, Hongyao, et al.
Published: (2025)
by: Tang, Hongyao, et al.
Published: (2025)
Stable Deep Reinforcement Learning via Isotropic Gaussian Representations
by: Pasand, Ali Saheb, et al.
Published: (2026)
by: Pasand, Ali Saheb, et al.
Published: (2026)
Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents
by: Obando-Ceron, Johan, et al.
Published: (2025)
by: Obando-Ceron, Johan, et al.
Published: (2025)
Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
by: Liu, Zihe, et al.
Published: (2025)
by: Liu, Zihe, et al.
Published: (2025)
Adaptive Computation Pruning for the Forgetting Transformer
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
CE-RM: A Pointwise Generative Reward Model Optimized via Two-Stage Rollout and Unified Criteria
by: Hu, Xinyu, et al.
Published: (2026)
by: Hu, Xinyu, et al.
Published: (2026)
A Mechanistic Analysis of Looped Reasoning Language Models
by: Blayney, Hugh, et al.
Published: (2026)
by: Blayney, Hugh, et al.
Published: (2026)
Stable Gradients for Stable Learning at Scale in Deep Reinforcement Learning
by: Castanyer, Roger Creus, et al.
Published: (2025)
by: Castanyer, Roger Creus, et al.
Published: (2025)
Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
Mixture of Experts in a Mixture of RL settings
by: Willi, Timon, et al.
Published: (2024)
by: Willi, Timon, et al.
Published: (2024)
Think-J: Learning to Think for Generative LLM-as-a-Judge
by: Huang, Hui, et al.
Published: (2025)
by: Huang, Hui, et al.
Published: (2025)
Learning Intractable Multimodal Policies with Reparameterization and Diversity Regularization
by: Wang, Ziqi, et al.
Published: (2025)
by: Wang, Ziqi, et al.
Published: (2025)
Complementary Reinforcement Learning
by: Muhtar, Dilxat, et al.
Published: (2026)
by: Muhtar, Dilxat, et al.
Published: (2026)
2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision
by: Li, Shilong, et al.
Published: (2024)
by: Li, Shilong, et al.
Published: (2024)
A Comedy of Estimators: On KL Regularization in RL Training of LLMs
by: Shah, Vedant, et al.
Published: (2025)
by: Shah, Vedant, et al.
Published: (2025)
Compositional Discrete Latent Code for High Fidelity, Productive Diffusion Models
by: Lavoie, Samuel, et al.
Published: (2025)
by: Lavoie, Samuel, et al.
Published: (2025)
ProgCo: Program Helps Self-Correction of Large Language Models
by: Song, Xiaoshuai, et al.
Published: (2025)
by: Song, Xiaoshuai, et al.
Published: (2025)
Versatile Energy-Based Probabilistic Models for High Energy Physics
by: Cheng, Taoli, et al.
Published: (2023)
by: Cheng, Taoli, et al.
Published: (2023)
BARD: budget-aware reasoning distillation
by: Niu, Lujie, et al.
Published: (2025)
by: Niu, Lujie, et al.
Published: (2025)
Distributional GFlowNets with Quantile Flows
by: Zhang, Dinghuai, et al.
Published: (2023)
by: Zhang, Dinghuai, et al.
Published: (2023)
Deconstructing Long Chain-of-Thought: A Structured Reasoning Optimization Framework for Long CoT Distillation
by: Luo, Yijia, et al.
Published: (2025)
by: Luo, Yijia, et al.
Published: (2025)
Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?
by: He, Yancheng, et al.
Published: (2025)
by: He, Yancheng, et al.
Published: (2025)
Not All LLM Reasoners Are Created Equal
by: Hosseini, Arian, et al.
Published: (2024)
by: Hosseini, Arian, et al.
Published: (2024)
On-Policy Optimization of ANFIS Policies Using Proximal Policy Optimization
by: Shankar, Kaaustaaub, et al.
Published: (2025)
by: Shankar, Kaaustaaub, et al.
Published: (2025)
Reparameterization Proximal Policy Optimization
by: Zhong, Hai, et al.
Published: (2025)
by: Zhong, Hai, et al.
Published: (2025)
Truncated Proximal Policy Optimization
by: Fan, Tiantian, et al.
Published: (2025)
by: Fan, Tiantian, et al.
Published: (2025)
Mixtures of Experts Unlock Parameter Scaling for Deep RL
by: Obando-Ceron, Johan, et al.
Published: (2024)
by: Obando-Ceron, Johan, et al.
Published: (2024)
Towards Sustainable Investment Policies Informed by Opponent Shaping
by: Duque, Juan Agustin, et al.
Published: (2026)
by: Duque, Juan Agustin, et al.
Published: (2026)
Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony
by: Lu, Han, et al.
Published: (2025)
by: Lu, Han, et al.
Published: (2025)
Managing multiple agents by automatically adjusting incentives
by: Akatsuka, Shunichi, et al.
Published: (2024)
by: Akatsuka, Shunichi, et al.
Published: (2024)
Why Open Source? A Game-Theoretic Analysis of the AI Race
by: Mladenovic, Andjela, et al.
Published: (2026)
by: Mladenovic, Andjela, et al.
Published: (2026)
Complexity-Regularized Proximal Policy Optimization
by: Serfilippi, Luca, et al.
Published: (2025)
by: Serfilippi, Luca, et al.
Published: (2025)
Beyond the Boundaries of Proximal Policy Optimization
by: Tan, Charlie B., et al.
Published: (2024)
by: Tan, Charlie B., et al.
Published: (2024)
Similar Items
-
The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning
by: Liu, Jiashun, et al.
Published: (2025) -
Neuroplastic Expansion in Deep Reinforcement Learning
by: Liu, Jiashun, et al.
Published: (2024) -
Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learning
by: Liu, Jiashun, et al.
Published: (2025) -
In value-based deep reinforcement learning, a pruned network is a good network
by: Obando-Ceron, Johan, et al.
Published: (2024) -
The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning Networks
by: Mayor, Walter, et al.
Published: (2025)