BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Guoxin, Meng, Fanzhe, Zhao, Jiale, Li, Minghao, Cheng, Daixuan, Song, Huatong, Chen, Jie, Lin, Yuzhi, Chen, Hui, Zhao, Xin, Song, Ruihua, Liu, Chang, Chen, Cheng, Jia, Kai, Wen, Ji-Rong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916048663478272
author Chen, Guoxin
Meng, Fanzhe
Zhao, Jiale
Li, Minghao
Cheng, Daixuan
Song, Huatong
Chen, Jie
Lin, Yuzhi
Chen, Hui
Zhao, Xin
Song, Ruihua
Liu, Chang
Chen, Cheng
Jia, Kai
Wen, Ji-Rong
author_facet Chen, Guoxin
Meng, Fanzhe
Zhao, Jiale
Li, Minghao
Cheng, Daixuan
Song, Huatong
Chen, Jie
Lin, Yuzhi
Chen, Hui
Zhao, Xin
Song, Ruihua
Liu, Chang
Chen, Cheng
Jia, Kai
Wen, Ji-Rong
contents Current code-agent benchmarks primarily evaluate localized issue resolution within a single target repository, leaving under-tested many software engineering tasks that require external knowledge or broader repository-level changes. We introduce BeyondSWE, a 500-instance benchmark drawn from 246 real-world GitHub repositories to evaluate code agents beyond single-repository bug fixing. BeyondSWE covers four representative settings: cross-repository issue resolution, domain-specific issue resolution, dependency-driven migration, and document-to-repository generation, spanning both broader knowledge scope and broader resolution scope. Our evaluation shows that BeyondSWE remains far from saturated: the best OpenHands-based agent reaches 46.12 average score, while the strongest Codex harness with GPT-5.4 (xhigh) reaches 56.65 under a search-aware prompt. To study whether external information access closes this gap, we use SearchSWE as a controlled diagnostic baseline for search-augmented coding. Search access improves most models and substantially helps some tasks, but the gains remain limited and uneven, showing that current agents still struggle to convert retrieved information into precise, version-compatible, and locally actionable code changes. These results suggest that deep search for coding remains an open problem: progress requires agents that can reliably combine external evidence with repository-local reasoning and execution-based verification.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03194
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?
Chen, Guoxin
Meng, Fanzhe
Zhao, Jiale
Li, Minghao
Cheng, Daixuan
Song, Huatong
Chen, Jie
Lin, Yuzhi
Chen, Hui
Zhao, Xin
Song, Ruihua
Liu, Chang
Chen, Cheng
Jia, Kai
Wen, Ji-Rong
Computation and Language
Software Engineering
Current code-agent benchmarks primarily evaluate localized issue resolution within a single target repository, leaving under-tested many software engineering tasks that require external knowledge or broader repository-level changes. We introduce BeyondSWE, a 500-instance benchmark drawn from 246 real-world GitHub repositories to evaluate code agents beyond single-repository bug fixing. BeyondSWE covers four representative settings: cross-repository issue resolution, domain-specific issue resolution, dependency-driven migration, and document-to-repository generation, spanning both broader knowledge scope and broader resolution scope. Our evaluation shows that BeyondSWE remains far from saturated: the best OpenHands-based agent reaches 46.12 average score, while the strongest Codex harness with GPT-5.4 (xhigh) reaches 56.65 under a search-aware prompt. To study whether external information access closes this gap, we use SearchSWE as a controlled diagnostic baseline for search-augmented coding. Search access improves most models and substantially helps some tasks, but the gains remain limited and uneven, showing that current agents still struggle to convert retrieved information into precise, version-compatible, and locally actionable code changes. These results suggest that deep search for coding remains an open problem: progress requires agents that can reliably combine external evidence with repository-local reasoning and execution-based verification.
title BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?
topic Computation and Language
Software Engineering
url https://arxiv.org/abs/2603.03194