Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Kunhao, Chambon, Pierre, Decugis, Juliette, Gehring, Jonas, Cohen, Taco, Negrevergne, Benjamin, Synnaeve, Gabriel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Makes Large Language Models Reason in (Multi-Turn) Code Generation?
by: Zheng, Kunhao, et al.
Published: (2024)
by: Zheng, Kunhao, et al.
Published: (2024)
RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
by: Gehring, Jonas, et al.
Published: (2024)
by: Gehring, Jonas, et al.
Published: (2024)
The KoLMogorov Test: Compression by Code Generation
by: Yoran, Ori, et al.
Published: (2025)
by: Yoran, Ori, et al.
Published: (2025)
BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
by: Chambon, Pierre, et al.
Published: (2025)
by: Chambon, Pierre, et al.
Published: (2025)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
by: Cohen, Taco, et al.
Published: (2025)
by: Cohen, Taco, et al.
Published: (2025)
A Deep Dive into Scaling RL for Code Generation with Synthetic Data and Curricula
by: Sancaktar, Cansu, et al.
Published: (2026)
by: Sancaktar, Cansu, et al.
Published: (2026)
Toward Training Superintelligent Software Agents through Self-Play SWE-RL
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
Meta Large Language Model Compiler: Foundation Models of Compiler Optimization
by: Cummins, Chris, et al.
Published: (2024)
by: Cummins, Chris, et al.
Published: (2024)
The Extrapolation Power of Implicit Models
by: Decugis, Juliette, et al.
Published: (2024)
by: Decugis, Juliette, et al.
Published: (2024)
Towards a Neural Debugger for Python
by: Beck, Maximilian, et al.
Published: (2026)
by: Beck, Maximilian, et al.
Published: (2026)
Improving Diversity in Language Models: When Temperature Fails, Change the Loss
by: Verine, Alexandre, et al.
Published: (2025)
by: Verine, Alexandre, et al.
Published: (2025)
Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
by: Hassid, Michael, et al.
Published: (2025)
by: Hassid, Michael, et al.
Published: (2025)
CWM: An Open-Weights LLM for Research on Code Generation with World Models
by: FAIR CodeGen team, et al.
Published: (2025)
by: FAIR CodeGen team, et al.
Published: (2025)
The Larger the Better? Improved LLM Code-Generation via Budget Reallocation
by: Hassid, Michael, et al.
Published: (2024)
by: Hassid, Michael, et al.
Published: (2024)
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
by: Butt, Natasha, et al.
Published: (2024)
by: Butt, Natasha, et al.
Published: (2024)
WybeCoder: Verified Imperative Code Generation
by: Gloeckle, Fabian, et al.
Published: (2026)
by: Gloeckle, Fabian, et al.
Published: (2026)
Extrapolation Merging: Keep Improving With Extrapolation and Merging
by: Lin, Yiguan, et al.
Published: (2025)
by: Lin, Yiguan, et al.
Published: (2025)
Beyond pass@k: Redundancy-Aware RLVR for Multi-Sample Code Generation
by: Florian, Le Bronnec, et al.
Published: (2026)
by: Florian, Le Bronnec, et al.
Published: (2026)
Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RL
by: Hong, Joey, et al.
Published: (2025)
by: Hong, Joey, et al.
Published: (2025)
Self-Execution Simulation Improves Coding Models
by: Maimon, Gallil, et al.
Published: (2026)
by: Maimon, Gallil, et al.
Published: (2026)
Model Extrapolation Expedites Alignment
by: Zheng, Chujie, et al.
Published: (2024)
by: Zheng, Chujie, et al.
Published: (2024)
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
by: Skorobogat, Ronald, et al.
Published: (2026)
by: Skorobogat, Ronald, et al.
Published: (2026)
WARM: On the Benefits of Weight Averaged Reward Models
by: Ramé, Alexandre, et al.
Published: (2024)
by: Ramé, Alexandre, et al.
Published: (2024)
ECCO: Can We Improve Model-Generated Code Efficiency Without Sacrificing Functional Correctness?
by: Waghjale, Siddhant, et al.
Published: (2024)
by: Waghjale, Siddhant, et al.
Published: (2024)
ContextRL: Enhancing MLLM's Knowledge Discovery Efficiency with Context-Augmented RL
by: Lu, Xingyu, et al.
Published: (2026)
by: Lu, Xingyu, et al.
Published: (2026)
Optimizing Language Models for Inference Time Objectives using Reinforcement Learning
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Mining the Mind: What 100M Beliefs Reveal About Frontier LLM Knowledge
by: Ghosh, Shrestha, et al.
Published: (2025)
by: Ghosh, Shrestha, et al.
Published: (2025)
Monte Carlo Graph Coloring
by: Cazenave, Tristan, et al.
Published: (2025)
by: Cazenave, Tristan, et al.
Published: (2025)
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
by: Jiang, Xue, et al.
Published: (2025)
by: Jiang, Xue, et al.
Published: (2025)
Incomplete Tasks Induce Shutdown Resistance in Some Frontier LLMs
by: Schlatter, Jeremy, et al.
Published: (2025)
by: Schlatter, Jeremy, et al.
Published: (2025)
Scaling Laws of RoPE-based Extrapolation
by: Liu, Xiaoran, et al.
Published: (2023)
by: Liu, Xiaoran, et al.
Published: (2023)
Extrapolation by Association: Length Generalization Transfer in Transformers
by: Cai, Ziyang, et al.
Published: (2025)
by: Cai, Ziyang, et al.
Published: (2025)
Getting the most out of your tokenizer for pre-training and domain adaptation
by: Dagan, Gautier, et al.
Published: (2024)
by: Dagan, Gautier, et al.
Published: (2024)
Rethinking Code Refinement: Learning to Judge Code Efficiency
by: Seo, Minju, et al.
Published: (2024)
by: Seo, Minju, et al.
Published: (2024)
Large Language Models for Extrapolative Modeling of Manufacturing Processes
by: Khanghah, Kiarash Naghavi, et al.
Published: (2025)
by: Khanghah, Kiarash Naghavi, et al.
Published: (2025)
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning
by: Perin, Gabriel J., et al.
Published: (2025)
by: Perin, Gabriel J., et al.
Published: (2025)
From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging
by: Shi, Yuling, et al.
Published: (2024)
by: Shi, Yuling, et al.
Published: (2024)
Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
by: Du, Weihua, et al.
Published: (2025)
by: Du, Weihua, et al.
Published: (2025)
Scaling Test-Time Compute for Agentic Coding
by: Kim, Joongwon, et al.
Published: (2026)
by: Kim, Joongwon, et al.
Published: (2026)
Similar Items
-
What Makes Large Language Models Reason in (Multi-Turn) Code Generation?
by: Zheng, Kunhao, et al.
Published: (2024) -
RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
by: Gehring, Jonas, et al.
Published: (2024) -
The KoLMogorov Test: Compression by Code Generation
by: Yoran, Ori, et al.
Published: (2025) -
BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
by: Chambon, Pierre, et al.
Published: (2025) -
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
by: Cohen, Taco, et al.
Published: (2025)