Agentic Reinforcement Learning for Real-World Code Repair

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Siyu, Karpovich, Anastasiya, Chen, Albert, Koscheka, Jessica, Jannu, Shailesh, Wen, Di, Zhu, Yuqing, Jain, Rohit, Geramifard, Alborz
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917042042437632
author Zhu, Siyu
Karpovich, Anastasiya
Chen, Albert
Koscheka, Jessica
Jannu, Shailesh
Wen, Di
Zhu, Yuqing
Jain, Rohit
Geramifard, Alborz
author_facet Zhu, Siyu
Karpovich, Anastasiya
Chen, Albert
Koscheka, Jessica
Jannu, Shailesh
Wen, Di
Zhu, Yuqing
Jain, Rohit
Geramifard, Alborz
contents We tackle the challenge of training reliable code-fixing agents in real repositories, where complex builds and shifting dependencies make evaluation unstable. We developed a verifiable pipeline with success defined as post-fix build validation and improved reproducibility across ~1K real issues by pinning dependencies and disabling automatic upgrades. Building on this, we introduced a scalable simplified pipeline for large-scale reinforcement learning (RL). Using this setup, we supervised fine-tuned Qwen3-32B in the full pipeline and applied RL on top of the SFT model in the simplified environment. The SFT model distilled from GPT-4.1 trajectories performs on par while being 56x smaller, and RL added 7-20% absolute gains under matched train-test conditions. "Thinking mode" was on par or worse in our experiments. Both SFT and RL models failed to generalize across environments, highlighting the importance of matching train-test environments for building reliable real-world code-fixing agents.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22075
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Agentic Reinforcement Learning for Real-World Code Repair
Zhu, Siyu
Karpovich, Anastasiya
Chen, Albert
Koscheka, Jessica
Jannu, Shailesh
Wen, Di
Zhu, Yuqing
Jain, Rohit
Geramifard, Alborz
Machine Learning
Artificial Intelligence
Computation and Language
We tackle the challenge of training reliable code-fixing agents in real repositories, where complex builds and shifting dependencies make evaluation unstable. We developed a verifiable pipeline with success defined as post-fix build validation and improved reproducibility across ~1K real issues by pinning dependencies and disabling automatic upgrades. Building on this, we introduced a scalable simplified pipeline for large-scale reinforcement learning (RL). Using this setup, we supervised fine-tuned Qwen3-32B in the full pipeline and applied RL on top of the SFT model in the simplified environment. The SFT model distilled from GPT-4.1 trajectories performs on par while being 56x smaller, and RL added 7-20% absolute gains under matched train-test conditions. "Thinking mode" was on par or worse in our experiments. Both SFT and RL models failed to generalize across environments, highlighting the importance of matching train-test environments for building reliable real-world code-fixing agents.
title Agentic Reinforcement Learning for Real-World Code Repair
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.22075