SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Soni, Aditya Bharat, Ghosh, Rajat, Bhargava, Vaishnavi, Chen, Valerie, Dutta, Debojyoti
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909995620106240
author Soni, Aditya Bharat
Ghosh, Rajat
Bhargava, Vaishnavi
Chen, Valerie
Dutta, Debojyoti
author_facet Soni, Aditya Bharat
Ghosh, Rajat
Bhargava, Vaishnavi
Chen, Valerie
Dutta, Debojyoti
contents Software testing is crucial for ensuring the correctness and reliability of software systems. Automated generation of issue reproduction tests from natural language issue descriptions enhances developer productivity by simplifying root cause analysis, promotes test-driven development -- "test first, write code later", and can be used for improving the effectiveness of automated issue resolution systems like coding agents. Existing methods proposed for this task predominantly rely on closed-source LLMs, with limited exploration of open models. To address this, we propose SWE-Tester -- a novel pipeline for training open-source LLMs to generate issue reproduction tests. First, we curate a high-quality training dataset of 41K instances from 2.6K open-source GitHub repositories and use it to train LLMs of varying sizes and families. The fine-tuned models achieve absolute improvements of up to 10\% in success rate and 21\% in change coverage on SWT-Bench Verified. Further analysis shows consistent improvements with increased inference-time compute, more data, and larger models. These results highlight the effectiveness of our framework for advancing open-source LLMs in this domain.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13713
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories
Soni, Aditya Bharat
Ghosh, Rajat
Bhargava, Vaishnavi
Chen, Valerie
Dutta, Debojyoti
Software Engineering
Machine Learning
Software testing is crucial for ensuring the correctness and reliability of software systems. Automated generation of issue reproduction tests from natural language issue descriptions enhances developer productivity by simplifying root cause analysis, promotes test-driven development -- "test first, write code later", and can be used for improving the effectiveness of automated issue resolution systems like coding agents. Existing methods proposed for this task predominantly rely on closed-source LLMs, with limited exploration of open models. To address this, we propose SWE-Tester -- a novel pipeline for training open-source LLMs to generate issue reproduction tests. First, we curate a high-quality training dataset of 41K instances from 2.6K open-source GitHub repositories and use it to train LLMs of varying sizes and families. The fine-tuned models achieve absolute improvements of up to 10\% in success rate and 21\% in change coverage on SWT-Bench Verified. Further analysis shows consistent improvements with increased inference-time compute, more data, and larger models. These results highlight the effectiveness of our framework for advancing open-source LLMs in this domain.
title SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories
topic Software Engineering
Machine Learning
url https://arxiv.org/abs/2601.13713