SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Soni, Aditya Bharat, Ghosh, Rajat, Bhargava, Vaishnavi, Chen, Valerie, Dutta, Debojyoti |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
CPP-UT-Bench: Can LLMs Write Complex Unit Tests in C++?
par: Bhargava, Vaishnavi, et autres
Publié: (2024)
par: Bhargava, Vaishnavi, et autres
Publié: (2024)
RANGER -- Repository-Level Agent for Graph-Enhanced Retrieval
par: Shah, Pratik, et autres
Publié: (2025)
par: Shah, Pratik, et autres
Publié: (2025)
A Multi-Agent Framework for Stateful Inference-Time Search
par: Lalan, Arshika, et autres
Publié: (2025)
par: Lalan, Arshika, et autres
Publié: (2025)
CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents
par: Pereira, Kristen, et autres
Publié: (2026)
par: Pereira, Kristen, et autres
Publié: (2026)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
par: Xu, Yisen, et autres
Publié: (2026)
par: Xu, Yisen, et autres
Publié: (2026)
SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?
par: He, Xinyi, et autres
Publié: (2025)
par: He, Xinyi, et autres
Publié: (2025)
SWE-Bench++: A Framework for the Scalable Generation of Software Engineering Benchmarks from Open-Source Repositories
par: Wang, Lilin, et autres
Publié: (2025)
par: Wang, Lilin, et autres
Publié: (2025)
SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
par: Miserendino, Samuel, et autres
Publié: (2025)
par: Miserendino, Samuel, et autres
Publié: (2025)
Reproduction Test Generation for Java SWE Issues
par: Ahmed, Toufique, et autres
Publié: (2026)
par: Ahmed, Toufique, et autres
Publié: (2026)
SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?
par: Ma, Jeffrey Jian, et autres
Publié: (2025)
par: Ma, Jeffrey Jian, et autres
Publié: (2025)
SWE-Mirror: Scaling Issue-Resolving Datasets by Mirroring Issues Across Repositories
par: Wang, Junhao, et autres
Publié: (2025)
par: Wang, Junhao, et autres
Publié: (2025)
SWE-Spot: Building Small Repo-Experts with Repository-Centric Learning
par: Peng, Jinjun, et autres
Publié: (2026)
par: Peng, Jinjun, et autres
Publié: (2026)
Predictive Scaling Laws for Efficient GRPO Training of Large Reasoning Models
par: Nimmaturi, Datta, et autres
Publié: (2025)
par: Nimmaturi, Datta, et autres
Publié: (2025)
Otter: Generating Tests from Issues to Validate SWE Patches
par: Ahmed, Toufique, et autres
Publié: (2025)
par: Ahmed, Toufique, et autres
Publié: (2025)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
par: Raghavendra, Mohit, et autres
Publié: (2026)
par: Raghavendra, Mohit, et autres
Publié: (2026)
Exploring the Lifecycle and Maintenance Practices of Pre-Trained Models in Open-Source Software Repositories
par: Koohjani, Matin, et autres
Publié: (2025)
par: Koohjani, Matin, et autres
Publié: (2025)
SWE-Exp: Experience-Driven Software Issue Resolution
par: Chen, Silin, et autres
Publié: (2025)
par: Chen, Silin, et autres
Publié: (2025)
Resolving Java Code Repository Issues with iSWE Agent
par: Ganhotra, Jatin, et autres
Publié: (2026)
par: Ganhotra, Jatin, et autres
Publié: (2026)
Automated Mapping of Vulnerability Advisories onto their Fix Commits in Open Source Repositories
par: Hommersom, Daan, et autres
Publié: (2021)
par: Hommersom, Daan, et autres
Publié: (2021)
An Empirical Validation of Open Source Repository Stability Metrics
par: Adejumo, Elijah Kayode, et autres
Publié: (2025)
par: Adejumo, Elijah Kayode, et autres
Publié: (2025)
Action Shapley: A Training Data Selection Metric for World Model in Reinforcement Learning
par: Ghosh, Rajat, et autres
Publié: (2026)
par: Ghosh, Rajat, et autres
Publié: (2026)
SWE-Universe: Scale Real-World Verifiable Environments to Millions
par: Chen, Mouxiang, et autres
Publié: (2026)
par: Chen, Mouxiang, et autres
Publié: (2026)
Investigating Test Overfitting on SWE-bench
par: Ahmed, Toufique, et autres
Publié: (2025)
par: Ahmed, Toufique, et autres
Publié: (2025)
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
par: Liang, Jiarong, et autres
Publié: (2026)
par: Liang, Jiarong, et autres
Publié: (2026)
GiveMeLabeledIssues: An Open Source Issue Recommendation System
par: Vargovich, Joseph, et autres
Publié: (2023)
par: Vargovich, Joseph, et autres
Publié: (2023)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
par: Aleithan, Reem, et autres
Publié: (2024)
par: Aleithan, Reem, et autres
Publié: (2024)
Breaking Single-Tester Limits: Multi-Agent LLMs for Multi-User Feature Testing
par: Feng, Sidong, et autres
Publié: (2025)
par: Feng, Sidong, et autres
Publié: (2025)
Classifying Issues in Open-source GitHub Repositories
par: Raaj, Amir Hossain, et autres
Publié: (2025)
par: Raaj, Amir Hossain, et autres
Publié: (2025)
BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
par: Zhou, Jinan, et autres
Publié: (2025)
par: Zhou, Jinan, et autres
Publié: (2025)
SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution
par: Li, Han, et autres
Publié: (2025)
par: Li, Han, et autres
Publié: (2025)
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
par: Cai, Songcheng, et autres
Publié: (2026)
par: Cai, Songcheng, et autres
Publié: (2026)
Revealing the value of Repository Centrality in lifespan prediction of Open Source Software Projects
par: He, Runzhi, et autres
Publié: (2024)
par: He, Runzhi, et autres
Publié: (2024)
The Product Beyond the Model -- An Empirical Study of Repositories of Open-Source ML Products
par: Nahar, Nadia, et autres
Publié: (2023)
par: Nahar, Nadia, et autres
Publié: (2023)
Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go
par: Pipalani, Yashshi, et autres
Publié: (2025)
par: Pipalani, Yashshi, et autres
Publié: (2025)
SWE-Arena: An Interactive Platform for Evaluating Foundation Models in Software Engineering
par: Zhao, Zhimin
Publié: (2025)
par: Zhao, Zhimin
Publié: (2025)
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
par: Jimenez, Carlos E., et autres
Publié: (2023)
par: Jimenez, Carlos E., et autres
Publié: (2023)
Are Autonomous Web Agents Good Testers?
par: Chevrot, Antoine, et autres
Publié: (2025)
par: Chevrot, Antoine, et autres
Publié: (2025)
HE-SNR: Uncovering Latent Logic via Entropy for Guiding Mid-Training on SWE-bench
par: Wang, Yueyang, et autres
Publié: (2026)
par: Wang, Yueyang, et autres
Publié: (2026)
R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
par: Jain, Naman, et autres
Publié: (2025)
par: Jain, Naman, et autres
Publié: (2025)
Can Old Tests Do New Tricks for Resolving SWE Issues?
par: Chen, Yang, et autres
Publié: (2025)
par: Chen, Yang, et autres
Publié: (2025)
Documents similaires
-
CPP-UT-Bench: Can LLMs Write Complex Unit Tests in C++?
par: Bhargava, Vaishnavi, et autres
Publié: (2024) -
RANGER -- Repository-Level Agent for Graph-Enhanced Retrieval
par: Shah, Pratik, et autres
Publié: (2025) -
A Multi-Agent Framework for Stateful Inference-Time Search
par: Lalan, Arshika, et autres
Publié: (2025) -
CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents
par: Pereira, Kristen, et autres
Publié: (2026) -
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
par: Xu, Yisen, et autres
Publié: (2026)