Breaking Training Bottlenecks: Effective and Stable Reinforcement Learning for Coding Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Zongqian, Huang, Shaohan, Chi, Zewen, Su, Yixuan, Zhou, Lexin, Dong, Li, Collier, Nigel, Wei, Furu
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918379502174208
author Li, Zongqian
Huang, Shaohan
Chi, Zewen
Su, Yixuan
Zhou, Lexin
Dong, Li
Collier, Nigel
Wei, Furu
author_facet Li, Zongqian
Huang, Shaohan
Chi, Zewen
Su, Yixuan
Zhou, Lexin
Dong, Li
Collier, Nigel
Wei, Furu
contents Modern code generation models exhibit longer outputs, accelerated capability growth, and changed training dynamics, rendering traditional training methodologies, algorithms, and datasets ineffective for improving their performance. To address these training bottlenecks, we propose MicroCoder-GRPO, an improved Group Relative Policy Optimization approach with three innovations: conditional truncation masking to improve long output potential while maintaining training stability, diversity-determined temperature selection to maintain and encourage output diversity, and removal of KL loss with high clipping ratios to facilitate solution diversity. MicroCoder-GRPO achieves up to 17.6% relative improvement over strong baselines on LiveCodeBench v6, with more pronounced gains under extended context evaluation. Additionally, we release MicroCoder-Dataset, a more challenging training corpus that achieves 3x larger performance gains than mainstream datasets on LiveCodeBench v6 within 300 training steps, and MicroCoder-Evaluator, a robust framework with approximately 25% improved evaluation accuracy and around 40% faster execution. Through comprehensive analysis across more than thirty controlled experiments, we reveal 34 training insights across seven main aspects, demonstrating that properly trained models can achieve competitive performance with larger counterparts.
format Preprint
id arxiv_https___arxiv_org_abs_2603_07777
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Breaking Training Bottlenecks: Effective and Stable Reinforcement Learning for Coding Models
Li, Zongqian
Huang, Shaohan
Chi, Zewen
Su, Yixuan
Zhou, Lexin
Dong, Li
Collier, Nigel
Wei, Furu
Machine Learning
Computation and Language
General Literature
Modern code generation models exhibit longer outputs, accelerated capability growth, and changed training dynamics, rendering traditional training methodologies, algorithms, and datasets ineffective for improving their performance. To address these training bottlenecks, we propose MicroCoder-GRPO, an improved Group Relative Policy Optimization approach with three innovations: conditional truncation masking to improve long output potential while maintaining training stability, diversity-determined temperature selection to maintain and encourage output diversity, and removal of KL loss with high clipping ratios to facilitate solution diversity. MicroCoder-GRPO achieves up to 17.6% relative improvement over strong baselines on LiveCodeBench v6, with more pronounced gains under extended context evaluation. Additionally, we release MicroCoder-Dataset, a more challenging training corpus that achieves 3x larger performance gains than mainstream datasets on LiveCodeBench v6 within 300 training steps, and MicroCoder-Evaluator, a robust framework with approximately 25% improved evaluation accuracy and around 40% faster execution. Through comprehensive analysis across more than thirty controlled experiments, we reveal 34 training insights across seven main aspects, demonstrating that properly trained models can achieve competitive performance with larger counterparts.
title Breaking Training Bottlenecks: Effective and Stable Reinforcement Learning for Coding Models
topic Machine Learning
Computation and Language
General Literature
url https://arxiv.org/abs/2603.07777