Agentic Reinforcement Learning for Real-World Code Repair
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917042042437632 |
|---|---|
| author | Zhu, Siyu Karpovich, Anastasiya Chen, Albert Koscheka, Jessica Jannu, Shailesh Wen, Di Zhu, Yuqing Jain, Rohit Geramifard, Alborz |
| author_facet | Zhu, Siyu Karpovich, Anastasiya Chen, Albert Koscheka, Jessica Jannu, Shailesh Wen, Di Zhu, Yuqing Jain, Rohit Geramifard, Alborz |
| contents | We tackle the challenge of training reliable code-fixing agents in real repositories, where complex builds and shifting dependencies make evaluation unstable. We developed a verifiable pipeline with success defined as post-fix build validation and improved reproducibility across ~1K real issues by pinning dependencies and disabling automatic upgrades. Building on this, we introduced a scalable simplified pipeline for large-scale reinforcement learning (RL). Using this setup, we supervised fine-tuned Qwen3-32B in the full pipeline and applied RL on top of the SFT model in the simplified environment. The SFT model distilled from GPT-4.1 trajectories performs on par while being 56x smaller, and RL added 7-20% absolute gains under matched train-test conditions. "Thinking mode" was on par or worse in our experiments. Both SFT and RL models failed to generalize across environments, highlighting the importance of matching train-test environments for building reliable real-world code-fixing agents. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_22075 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Agentic Reinforcement Learning for Real-World Code Repair Zhu, Siyu Karpovich, Anastasiya Chen, Albert Koscheka, Jessica Jannu, Shailesh Wen, Di Zhu, Yuqing Jain, Rohit Geramifard, Alborz Machine Learning Artificial Intelligence Computation and Language We tackle the challenge of training reliable code-fixing agents in real repositories, where complex builds and shifting dependencies make evaluation unstable. We developed a verifiable pipeline with success defined as post-fix build validation and improved reproducibility across ~1K real issues by pinning dependencies and disabling automatic upgrades. Building on this, we introduced a scalable simplified pipeline for large-scale reinforcement learning (RL). Using this setup, we supervised fine-tuned Qwen3-32B in the full pipeline and applied RL on top of the SFT model in the simplified environment. The SFT model distilled from GPT-4.1 trajectories performs on par while being 56x smaller, and RL added 7-20% absolute gains under matched train-test conditions. "Thinking mode" was on par or worse in our experiments. Both SFT and RL models failed to generalize across environments, highlighting the importance of matching train-test environments for building reliable real-world code-fixing agents. |
| title | Agentic Reinforcement Learning for Real-World Code Repair |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2510.22075 |