SWE-Spot: Building Small Repo-Experts with Repository-Centric Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Peng, Jinjun, Saebo, Magnus, Zhong, Tianjun, Cheng, Yi-Jie, Yang, Junfeng, Ray, Baishakhi, Chen, Simin, Ding, Yangruibo
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912859081932800
author Peng, Jinjun
Saebo, Magnus
Zhong, Tianjun
Cheng, Yi-Jie
Yang, Junfeng
Ray, Baishakhi
Chen, Simin
Ding, Yangruibo
author_facet Peng, Jinjun
Saebo, Magnus
Zhong, Tianjun
Cheng, Yi-Jie
Yang, Junfeng
Ray, Baishakhi
Chen, Simin
Ding, Yangruibo
contents The deployment of coding agents in privacy-sensitive and resource-constrained environments drives the demand for capable open-weight Small Language Models (SLMs). However, they suffer from a fundamental capability gap: unlike frontier large models, they lack the inference-time strong generalization to work with complicated, unfamiliar codebases. We identify that the prevailing Task-Centric Learning (TCL) paradigm, which scales exposure across disparate repositories, fails to address this limitation. In response, we propose Repository-Centric Learning (RCL), a paradigm shift that prioritizes vertical repository depth over horizontal task breadth, suggesting SLMs must internalize the "physics" of a target software environment through parametric knowledge acquisition, rather than attempting to recover it via costly inference-time search. Following this new paradigm, we design a four-unit Repository-Centric Experience, transforming static codebases into interactive learning signals, to train SWE-Spot-4B, a family of highly compact models built as repo-specialized experts that breaks established scaling trends, outperforming open-weight models up to larger (e.g., CWM by Meta, Qwen3-Coder-30B) and surpassing/matching efficiency-focused commercial models (e.g., GPT-4.1-mini, GPT-5-nano) across multiple SWE tasks. Further analysis reveals that RCL yields higher training sample efficiency and lower inference costs, emphasizing that for building efficient intelligence, repository mastery is a distinct and necessary dimension that complements general coding capability.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21649
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SWE-Spot: Building Small Repo-Experts with Repository-Centric Learning
Peng, Jinjun
Saebo, Magnus
Zhong, Tianjun
Cheng, Yi-Jie
Yang, Junfeng
Ray, Baishakhi
Chen, Simin
Ding, Yangruibo
Machine Learning
Artificial Intelligence
Computation and Language
Software Engineering
The deployment of coding agents in privacy-sensitive and resource-constrained environments drives the demand for capable open-weight Small Language Models (SLMs). However, they suffer from a fundamental capability gap: unlike frontier large models, they lack the inference-time strong generalization to work with complicated, unfamiliar codebases. We identify that the prevailing Task-Centric Learning (TCL) paradigm, which scales exposure across disparate repositories, fails to address this limitation. In response, we propose Repository-Centric Learning (RCL), a paradigm shift that prioritizes vertical repository depth over horizontal task breadth, suggesting SLMs must internalize the "physics" of a target software environment through parametric knowledge acquisition, rather than attempting to recover it via costly inference-time search. Following this new paradigm, we design a four-unit Repository-Centric Experience, transforming static codebases into interactive learning signals, to train SWE-Spot-4B, a family of highly compact models built as repo-specialized experts that breaks established scaling trends, outperforming open-weight models up to larger (e.g., CWM by Meta, Qwen3-Coder-30B) and surpassing/matching efficiency-focused commercial models (e.g., GPT-4.1-mini, GPT-5-nano) across multiple SWE tasks. Further analysis reveals that RCL yields higher training sample efficiency and lower inference costs, emphasizing that for building efficient intelligence, repository mastery is a distinct and necessary dimension that complements general coding capability.
title SWE-Spot: Building Small Repo-Experts with Repository-Centric Learning
topic Machine Learning
Artificial Intelligence
Computation and Language
Software Engineering
url https://arxiv.org/abs/2601.21649