Breaking Training Bottlenecks: Effective and Stable Reinforcement Learning for Coding Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zongqian, Huang, Shaohan, Chi, Zewen, Su, Yixuan, Zhou, Lexin, Dong, Li, Collier, Nigel, Wei, Furu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Data Difficulty: Improving Coding Models via Reinforcement Learning on Fresh and Challenging Problems
by: Li, Zongqian, et al.
Published: (2026)
by: Li, Zongqian, et al.
Published: (2026)
A Survey on Prompt Tuning
by: Li, Zongqian, et al.
Published: (2025)
by: Li, Zongqian, et al.
Published: (2025)
PT-MoE: An Efficient Finetuning Framework for Integrating Mixture-of-Experts into Prompt Tuning
by: Li, Zongqian, et al.
Published: (2025)
by: Li, Zongqian, et al.
Published: (2025)
500xCompressor: Generalized Prompt Compression for Large Language Models
by: Li, Zongqian, et al.
Published: (2024)
by: Li, Zongqian, et al.
Published: (2024)
Towards Stable and Effective Reinforcement Learning for Mixture-of-Experts
by: Zhang, Di, et al.
Published: (2025)
by: Zhang, Di, et al.
Published: (2025)
Prompt Compression for Large Language Models: A Survey
by: Li, Zongqian, et al.
Published: (2024)
by: Li, Zongqian, et al.
Published: (2024)
Flexi-LoRA with Input-Adaptive Ranks: Efficient Finetuning for Speech and Reasoning Tasks
by: Li, Zongqian, et al.
Published: (2026)
by: Li, Zongqian, et al.
Published: (2026)
A guideline for the methodology chapter in computer science dissertations
by: Araujo, Marco
Published: (2024)
by: Araujo, Marco
Published: (2024)
POSIM: A Multi-Agent Simulation Framework for Social Media Public Opinion Evolution and Governance
by: Zhang, Yongmao, et al.
Published: (2026)
by: Zhang, Yongmao, et al.
Published: (2026)
People, IT, and Structuration (PIS): An Integrative Theoretical Framework for Management Information Systems
by: Huang, Wei, et al.
Published: (2026)
by: Huang, Wei, et al.
Published: (2026)
Max Bense as a Visionary: from Entropy to the Dialectics of Programmed Images
by: Robillard, Gaëtan
Published: (2024)
by: Robillard, Gaëtan
Published: (2024)
A Brief History of Fréchet Distances: From Curves and Probability Laws to FID
by: Wu, Yuli
Published: (2026)
by: Wu, Yuli
Published: (2026)
Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology
by: Jiao, Yuchen, et al.
Published: (2025)
by: Jiao, Yuchen, et al.
Published: (2025)
Tensor Cookbook: Mastering Tensors through Diagrams
by: Rakhshan, Beheshteh T., et al.
Published: (2026)
by: Rakhshan, Beheshteh T., et al.
Published: (2026)
Reproducibility: The New Frontier in AI Governance
by: Mason-Williams, Israel, et al.
Published: (2025)
by: Mason-Williams, Israel, et al.
Published: (2025)
The history of digital ethics
by: Müller, Vincent C.
Published: (2025)
by: Müller, Vincent C.
Published: (2025)
The History of Quantum Games
by: Piispanen, Laura, et al.
Published: (2023)
by: Piispanen, Laura, et al.
Published: (2023)
Validation of a Small Language Model for DSM-5 Substance Category Classification in Child Welfare Records
by: Perron, Brian E., et al.
Published: (2026)
by: Perron, Brian E., et al.
Published: (2026)
Automation and AI Technology in Surface Mining With a Brief Introduction to Open-Pit Operations in the Pilbara
by: Leung, Raymond, et al.
Published: (2023)
by: Leung, Raymond, et al.
Published: (2023)
The Era of Agentic Organization: Learning to Organize with Language Models
by: Chi, Zewen, et al.
Published: (2025)
by: Chi, Zewen, et al.
Published: (2025)
Revista Chilena de Literatura
Published: (2005)
Published: (2005)
ReasonGraph: Visualisation of Reasoning Paths
by: Li, Zongqian, et al.
Published: (2025)
by: Li, Zongqian, et al.
Published: (2025)
Black-Box On-Policy Distillation of Large Language Models
by: Ye, Tianzhu, et al.
Published: (2025)
by: Ye, Tianzhu, et al.
Published: (2025)
On-Policy RL with Optimal Reward Baseline
by: Hao, Yaru, et al.
Published: (2025)
by: Hao, Yaru, et al.
Published: (2025)
Aksara
Published: (2017)
Published: (2017)
Manuscript and Print in Late Medieval and Early Modern Britain
Published: (2021)
Published: (2021)
Education Paradigm Shift To Maintain Human Competitive Advantage Over AI
by: Selitskiy, Stanislav, et al.
Published: (2025)
by: Selitskiy, Stanislav, et al.
Published: (2025)
Altre Modernità
Published: (2009)
Published: (2009)
Reward Reasoning Model
by: Guo, Jiaxin, et al.
Published: (2025)
by: Guo, Jiaxin, et al.
Published: (2025)
Darwin Turing Dawkins: Building a General Theory of Evolution
by: Adleman, Leonard M.
Published: (2024)
by: Adleman, Leonard M.
Published: (2024)
From Physical Difference to Meaning: A Constructor-Theoretic Framework for Prebiotic Information in Casimir-Lifshitz-Coupled Protocell Clusters
by: Massoth, Michael
Published: (2026)
by: Massoth, Michael
Published: (2026)
Kaggle Chronicles: 15 Years of Competitions, Community and Data Science Innovation
by: Bönisch, Kevin, et al.
Published: (2025)
by: Bönisch, Kevin, et al.
Published: (2025)
Similar Items
-
Scaling Data Difficulty: Improving Coding Models via Reinforcement Learning on Fresh and Challenging Problems
by: Li, Zongqian, et al.
Published: (2026) -
A Survey on Prompt Tuning
by: Li, Zongqian, et al.
Published: (2025) -
PT-MoE: An Efficient Finetuning Framework for Integrating Mixture-of-Experts into Prompt Tuning
by: Li, Zongqian, et al.
Published: (2025) -
500xCompressor: Generalized Prompt Compression for Large Language Models
by: Li, Zongqian, et al.
Published: (2024) -
Towards Stable and Effective Reinforcement Learning for Mixture-of-Experts
by: Zhang, Di, et al.
Published: (2025)