BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916048663478272 |
|---|---|
| author | Chen, Guoxin Meng, Fanzhe Zhao, Jiale Li, Minghao Cheng, Daixuan Song, Huatong Chen, Jie Lin, Yuzhi Chen, Hui Zhao, Xin Song, Ruihua Liu, Chang Chen, Cheng Jia, Kai Wen, Ji-Rong |
| author_facet | Chen, Guoxin Meng, Fanzhe Zhao, Jiale Li, Minghao Cheng, Daixuan Song, Huatong Chen, Jie Lin, Yuzhi Chen, Hui Zhao, Xin Song, Ruihua Liu, Chang Chen, Cheng Jia, Kai Wen, Ji-Rong |
| contents | Current code-agent benchmarks primarily evaluate localized issue resolution within a single target repository, leaving under-tested many software engineering tasks that require external knowledge or broader repository-level changes. We introduce BeyondSWE, a 500-instance benchmark drawn from 246 real-world GitHub repositories to evaluate code agents beyond single-repository bug fixing. BeyondSWE covers four representative settings: cross-repository issue resolution, domain-specific issue resolution, dependency-driven migration, and document-to-repository generation, spanning both broader knowledge scope and broader resolution scope. Our evaluation shows that BeyondSWE remains far from saturated: the best OpenHands-based agent reaches 46.12 average score, while the strongest Codex harness with GPT-5.4 (xhigh) reaches 56.65 under a search-aware prompt. To study whether external information access closes this gap, we use SearchSWE as a controlled diagnostic baseline for search-augmented coding. Search access improves most models and substantially helps some tasks, but the gains remain limited and uneven, showing that current agents still struggle to convert retrieved information into precise, version-compatible, and locally actionable code changes. These results suggest that deep search for coding remains an open problem: progress requires agents that can reliably combine external evidence with repository-local reasoning and execution-based verification. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_03194 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing? Chen, Guoxin Meng, Fanzhe Zhao, Jiale Li, Minghao Cheng, Daixuan Song, Huatong Chen, Jie Lin, Yuzhi Chen, Hui Zhao, Xin Song, Ruihua Liu, Chang Chen, Cheng Jia, Kai Wen, Ji-Rong Computation and Language Software Engineering Current code-agent benchmarks primarily evaluate localized issue resolution within a single target repository, leaving under-tested many software engineering tasks that require external knowledge or broader repository-level changes. We introduce BeyondSWE, a 500-instance benchmark drawn from 246 real-world GitHub repositories to evaluate code agents beyond single-repository bug fixing. BeyondSWE covers four representative settings: cross-repository issue resolution, domain-specific issue resolution, dependency-driven migration, and document-to-repository generation, spanning both broader knowledge scope and broader resolution scope. Our evaluation shows that BeyondSWE remains far from saturated: the best OpenHands-based agent reaches 46.12 average score, while the strongest Codex harness with GPT-5.4 (xhigh) reaches 56.65 under a search-aware prompt. To study whether external information access closes this gap, we use SearchSWE as a controlled diagnostic baseline for search-augmented coding. Search access improves most models and substantially helps some tasks, but the gains remain limited and uneven, showing that current agents still struggle to convert retrieved information into precise, version-compatible, and locally actionable code changes. These results suggest that deep search for coding remains an open problem: progress requires agents that can reliably combine external evidence with repository-local reasoning and execution-based verification. |
| title | BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing? |
| topic | Computation and Language Software Engineering |
| url | https://arxiv.org/abs/2603.03194 |