Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Zhenpeng, Pan, Leiyu, Bai, Xue, Liu, Dening, Dong, Guanting, Huang, Jiaming, Lv, Minxuan, Hu, Wenping, Zhang, Fuzheng, Gai, Kun, Zhou, Guorui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning
by: Su, Zhenpeng, et al.
Published: (2025)
by: Su, Zhenpeng, et al.
Published: (2025)
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning
by: Su, Zhenpeng, et al.
Published: (2025)
by: Su, Zhenpeng, et al.
Published: (2025)
Good Reasoning Makes Good Demonstrations: Implicit Reasoning Quality Supervision via In-Context Reinforcement Learning
by: Mei, Tiehua, et al.
Published: (2026)
by: Mei, Tiehua, et al.
Published: (2026)
Klear-AgentForge: Forging Agentic Intelligence through Posttraining Scaling
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Klear-CodeTest: Scalable Test Case Generation for Code Reinforcement Learning
by: Fu, Jia, et al.
Published: (2025)
by: Fu, Jia, et al.
Published: (2025)
Leanabell-Prover: Posttraining Scaling in Formal Reasoning
by: Zhang, Jingyuan, et al.
Published: (2025)
by: Zhang, Jingyuan, et al.
Published: (2025)
DSMoE: Matrix-Partitioned Experts with Dynamic Routing for Computation-Efficient Dense LLMs
by: Lv, Minxuan, et al.
Published: (2025)
by: Lv, Minxuan, et al.
Published: (2025)
Finedeep: Mitigating Sparse Activation in Dense LLMs via Multi-Layer Fine-Grained Experts
by: Pan, Leiyu, et al.
Published: (2025)
by: Pan, Leiyu, et al.
Published: (2025)
Agentic Entropy-Balanced Policy Optimization
by: Dong, Guanting, et al.
Published: (2025)
by: Dong, Guanting, et al.
Published: (2025)
Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR
by: Wang, Jiakang, et al.
Published: (2025)
by: Wang, Jiakang, et al.
Published: (2025)
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
by: Liu, Runze, et al.
Published: (2025)
by: Liu, Runze, et al.
Published: (2025)
When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL
by: Wang, Jiakang, et al.
Published: (2025)
by: Wang, Jiakang, et al.
Published: (2025)
Leanabell-Prover-V2: Verifier-integrated Reasoning for Formal Theorem Proving via Reinforcement Learning
by: Ji, Xingguang, et al.
Published: (2025)
by: Ji, Xingguang, et al.
Published: (2025)
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
by: Zhang, Hongzhi, et al.
Published: (2025)
by: Zhang, Hongzhi, et al.
Published: (2025)
Agentic Reinforced Policy Optimization
by: Dong, Guanting, et al.
Published: (2025)
by: Dong, Guanting, et al.
Published: (2025)
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
by: Lv, Minxuan, et al.
Published: (2026)
by: Lv, Minxuan, et al.
Published: (2026)
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
by: Liu, Xiao, et al.
Published: (2024)
by: Liu, Xiao, et al.
Published: (2024)
WebThinker: Empowering Large Reasoning Models with Deep Research Capability
by: Li, Xiaoxi, et al.
Published: (2025)
by: Li, Xiaoxi, et al.
Published: (2025)
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
by: Marshall, Noah, et al.
Published: (2024)
by: Marshall, Noah, et al.
Published: (2024)
DCPO: Dynamic Clipping Policy Optimization
by: Yang, Shihui, et al.
Published: (2025)
by: Yang, Shihui, et al.
Published: (2025)
Optimized Gradient Clipping for Noisy Label Learning
by: Ye, Xichen, et al.
Published: (2024)
by: Ye, Xichen, et al.
Published: (2024)
Just Ask One More Time! Self-Agreement Improves Reasoning of Language Models in (Almost) All Scenarios
by: Lin, Lei, et al.
Published: (2023)
by: Lin, Lei, et al.
Published: (2023)
State Regularized Policy Optimization on Data with Dynamics Shift
by: Xue, Zhenghai, et al.
Published: (2023)
by: Xue, Zhenghai, et al.
Published: (2023)
Robust Stochastic Optimization via Gradient Quantile Clipping
by: Merad, Ibrahim, et al.
Published: (2023)
by: Merad, Ibrahim, et al.
Published: (2023)
SetPO: Set-Level Policy Optimization for Diversity-Preserving LLM Reasoning
by: Li, Chenyi, et al.
Published: (2026)
by: Li, Chenyi, et al.
Published: (2026)
An Empirical Study on the Robustness of Massively Multilingual Neural Machine Translation
by: Supryadi, et al.
Published: (2024)
by: Supryadi, et al.
Published: (2024)
Reinforced Preference Optimization for Reasoning-Augmented Recommendations
by: Gao, Jingtong, et al.
Published: (2026)
by: Gao, Jingtong, et al.
Published: (2026)
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
by: Cheng, Junhao, et al.
Published: (2026)
by: Cheng, Junhao, et al.
Published: (2026)
Toward Effective Tool-Integrated Reasoning via Self-Evolved Preference Learning
by: Chen, Yifei, et al.
Published: (2025)
by: Chen, Yifei, et al.
Published: (2025)
Design and Development of Laughter Recognition System Based on Multimodal Fusion and Deep Learning
by: Zhao, Fuzheng, et al.
Published: (2024)
by: Zhao, Fuzheng, et al.
Published: (2024)
Beyond Token-Level Policy Gradients for Complex Reasoning with Large Language Models
by: Xu, Mufan, et al.
Published: (2026)
by: Xu, Mufan, et al.
Published: (2026)
Holistic Capability Preservation: Towards Compact Yet Comprehensive Reasoning Models
by: Ling Team, et al.
Published: (2025)
by: Ling Team, et al.
Published: (2025)
ET-Agent: Incentivizing Effective Tool-Integrated Reasoning Agent via Behavior Calibration
by: Chen, Yifei, et al.
Published: (2026)
by: Chen, Yifei, et al.
Published: (2026)
Inductive-Deductive Strategy Reuse for Multi-Turn Instructional Dialogues
by: Ou, Jiao, et al.
Published: (2024)
by: Ou, Jiao, et al.
Published: (2024)
ERABAL: Enhancing Role-Playing Agents through Boundary-Aware Learning
by: Tang, Yihong, et al.
Published: (2024)
by: Tang, Yihong, et al.
Published: (2024)
Enhancing Role-playing Systems through Aggressive Queries: Evaluation and Improvement
by: Tang, Yihong, et al.
Published: (2024)
by: Tang, Yihong, et al.
Published: (2024)
DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning
by: Li, Chengpeng, et al.
Published: (2024)
by: Li, Chengpeng, et al.
Published: (2024)
Clipping-Free Policy Optimization for Large Language Models
by: Çağatan, Ömer Veysel, et al.
Published: (2026)
by: Çağatan, Ömer Veysel, et al.
Published: (2026)
HoME: Hierarchy of Multi-Gate Experts for Multi-Task Learning at Kuaishou
by: Wang, Xu, et al.
Published: (2024)
by: Wang, Xu, et al.
Published: (2024)
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
by: Dong, Guanting, et al.
Published: (2025)
by: Dong, Guanting, et al.
Published: (2025)
Similar Items
-
CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning
by: Su, Zhenpeng, et al.
Published: (2025) -
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning
by: Su, Zhenpeng, et al.
Published: (2025) -
Good Reasoning Makes Good Demonstrations: Implicit Reasoning Quality Supervision via In-Context Reinforcement Learning
by: Mei, Tiehua, et al.
Published: (2026) -
Klear-AgentForge: Forging Agentic Intelligence through Posttraining Scaling
by: Wang, Qi, et al.
Published: (2025) -
Klear-CodeTest: Scalable Test Case Generation for Code Reinforcement Learning
by: Fu, Jia, et al.
Published: (2025)