SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Raghavendra, Mohit, Dan, Soham, Calvo, Miguel Romero, He, Yannis Yiming, Mols, Johannes Baptist, Anand, Gautam, McCollum, Cole, Arakelyan, Edgar, Bharadwaj, Vijay, Park, Andrew, Da, Jeff, Rezaei, MohammadHossein, Liu, Bing, Kenstler, Brad, He, Yunzhong
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913104280944640
author Raghavendra, Mohit
Dan, Soham
Calvo, Miguel Romero
He, Yannis Yiming
Mols, Johannes Baptist
Anand, Gautam
McCollum, Cole
Arakelyan, Edgar
Bharadwaj, Vijay
Park, Andrew
Da, Jeff
Rezaei, MohammadHossein
Liu, Bing
Kenstler, Brad
He, Yunzhong
author_facet Raghavendra, Mohit
Dan, Soham
Calvo, Miguel Romero
He, Yannis Yiming
Mols, Johannes Baptist
Anand, Gautam
McCollum, Cole
Arakelyan, Edgar
Bharadwaj, Vijay
Park, Andrew
Da, Jeff
Rezaei, MohammadHossein
Liu, Bing
Kenstler, Brad
He, Yunzhong
contents We introduce SWE Atlas, a benchmark suite for coding agents spanning three professional software engineering workflows: Codebase Q&A (124 tasks), Test Writing (90 tasks), and Refactoring (70 tasks). SWE Atlas differs from prior SWE benchmarks in three key ways: it targets underrepresented but practically important task categories, uses comprehensive category-specific evaluation protocols, and adopts under-specified, agentic task formulations that better reflect real-world usage. Its evaluation framework combines programmatic checks with rubric-based assessment. This goes beyond functional correctness, evaluating software engineering quality, including test and refactor completeness, maintainability, reusable abstractions, and codebase hygiene. We evaluate a range of frontier and open-weight models on SWE Atlas and find that GPT-5.4 and Opus 4.7 achieve the strongest overall performance, while even the best open-weight models score poorly. Our analysis suggests that top models rely on extensive codebase exploration and runtime-driven reasoning. However, even top models consistently struggle with subtle edge cases, complex runtime analysis, and adherence to software engineering best practices. Overall, SWE Atlas provides a complementary evaluation suite for measuring both correctness and engineering quality in coding agents.
format Preprint
id arxiv_https___arxiv_org_abs_2605_08366
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
Raghavendra, Mohit
Dan, Soham
Calvo, Miguel Romero
He, Yannis Yiming
Mols, Johannes Baptist
Anand, Gautam
McCollum, Cole
Arakelyan, Edgar
Bharadwaj, Vijay
Park, Andrew
Da, Jeff
Rezaei, MohammadHossein
Liu, Bing
Kenstler, Brad
He, Yunzhong
Machine Learning
Software Engineering
We introduce SWE Atlas, a benchmark suite for coding agents spanning three professional software engineering workflows: Codebase Q&A (124 tasks), Test Writing (90 tasks), and Refactoring (70 tasks). SWE Atlas differs from prior SWE benchmarks in three key ways: it targets underrepresented but practically important task categories, uses comprehensive category-specific evaluation protocols, and adopts under-specified, agentic task formulations that better reflect real-world usage. Its evaluation framework combines programmatic checks with rubric-based assessment. This goes beyond functional correctness, evaluating software engineering quality, including test and refactor completeness, maintainability, reusable abstractions, and codebase hygiene. We evaluate a range of frontier and open-weight models on SWE Atlas and find that GPT-5.4 and Opus 4.7 achieve the strongest overall performance, while even the best open-weight models score poorly. Our analysis suggests that top models rely on extensive codebase exploration and runtime-driven reasoning. However, even top models consistently struggle with subtle edge cases, complex runtime analysis, and adherence to software engineering best practices. Overall, SWE Atlas provides a complementary evaluation suite for measuring both correctness and engineering quality in coding agents.
title SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
topic Machine Learning
Software Engineering
url https://arxiv.org/abs/2605.08366