BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Guoxin, Meng, Fanzhe, Zhao, Jiale, Li, Minghao, Cheng, Daixuan, Song, Huatong, Chen, Jie, Lin, Yuzhi, Chen, Hui, Zhao, Xin, Song, Ruihua, Liu, Chang, Chen, Cheng, Jia, Kai, Wen, Ji-Rong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Toward Autonomous Long-Horizon Engineering for ML Research
by: Chen, Guoxin, et al.
Published: (2026)
by: Chen, Guoxin, et al.
Published: (2026)
SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training
by: Song, Huatong, et al.
Published: (2026)
by: Song, Huatong, et al.
Published: (2026)
Immersion in the GitHub Universe: Scaling Coding Agents to Mastery
by: Zhao, Jiale, et al.
Published: (2026)
by: Zhao, Jiale, et al.
Published: (2026)
Computer Environments Elicit General Agentic Intelligence in LLMs
by: Cheng, Daixuan, et al.
Published: (2026)
by: Cheng, Daixuan, et al.
Published: (2026)
SWE-Spot: Building Small Repo-Experts with Repository-Centric Learning
by: Peng, Jinjun, et al.
Published: (2026)
by: Peng, Jinjun, et al.
Published: (2026)
Beyond A Fixed Seal: Adaptive Stealing Watermark in Large Language Models
by: Zhang, Shuhao, et al.
Published: (2026)
by: Zhang, Shuhao, et al.
Published: (2026)
SWE-World: Building Software Engineering Agents in Docker-Free Environments
by: Sun, Shuang, et al.
Published: (2026)
by: Sun, Shuang, et al.
Published: (2026)
RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale
by: Chen, Zhilong, et al.
Published: (2025)
by: Chen, Zhilong, et al.
Published: (2025)
Coffee: Boost Your Code LLMs by Fixing Bugs with Feedback
by: Moon, Seungjun, et al.
Published: (2023)
by: Moon, Seungjun, et al.
Published: (2023)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
by: Raghavendra, Mohit, et al.
Published: (2026)
by: Raghavendra, Mohit, et al.
Published: (2026)
ReForm: Reflective Autoformalization with Prospective Bounded Sequence Optimization
by: Chen, Guoxin, et al.
Published: (2025)
by: Chen, Guoxin, et al.
Published: (2025)
SWE-Synth: Synthesizing Verifiable Bug-Fix Data to Enable Large Language Models in Resolving Real-World Bugs
by: Pham, Minh V. T., et al.
Published: (2025)
by: Pham, Minh V. T., et al.
Published: (2025)
RepoZero: Can LLMs Generate a Code Repository from Scratch?
by: Zhang, Zhaoxi, et al.
Published: (2026)
by: Zhang, Zhaoxi, et al.
Published: (2026)
Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents
by: Huang, Jiawei, et al.
Published: (2026)
by: Huang, Jiawei, et al.
Published: (2026)
R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
by: Song, Huatong, et al.
Published: (2025)
by: Song, Huatong, et al.
Published: (2025)
Debug2Fix: Can Interactive Debugging Help Coding Agents Fix More Bugs?
by: Garg, Spandan, et al.
Published: (2026)
by: Garg, Spandan, et al.
Published: (2026)
Artifacts of the ADS Bug-Fix Pattern Study
by: Chen, Yuntianyi, et al.
Published: (2025)
by: Chen, Yuntianyi, et al.
Published: (2025)
Beyond Bug Fixes: An Empirical Investigation of Post-Merge Code Quality Issues in Agent-Generated Pull Requests
by: Cynthia, Shamse Tasnim, et al.
Published: (2026)
by: Cynthia, Shamse Tasnim, et al.
Published: (2026)
MarsCode Agent: AI-native Automated Bug Fixing
by: Liu, Yizhou, et al.
Published: (2024)
by: Liu, Yizhou, et al.
Published: (2024)
BugPilot: Complex Bug Generation for Efficient Learning of SWE Skills
by: Sonwane, Atharv, et al.
Published: (2025)
by: Sonwane, Atharv, et al.
Published: (2025)
VeriGRAG: Enhancing LLM-Based Verilog Code Generation with Structure-Aware Soft Prompts
by: Zhao, Jiayu, et al.
Published: (2025)
by: Zhao, Jiayu, et al.
Published: (2025)
Pull Requests as a Training Signal for Repo-Level Code Editing
by: Zhu, Qinglin, et al.
Published: (2026)
by: Zhu, Qinglin, et al.
Published: (2026)
Can Students Beyond The Teacher? Distilling Knowledge from Teacher's Bias
by: Zhang, Jianhua, et al.
Published: (2024)
by: Zhang, Jianhua, et al.
Published: (2024)
Aletheia: Quantifying Cognitive Conviction in Reasoning Models via Regularized Inverse Confusion Matrix
by: Fu, Fanzhe
Published: (2026)
by: Fu, Fanzhe
Published: (2026)
The Meta-Prompting Protocol: Orchestrating LLMs via Adversarial Feedback Loops
by: Fu, Fanzhe
Published: (2025)
by: Fu, Fanzhe
Published: (2025)
BugsRepo: A Comprehensive Curated Dataset of Bug Reports, Comments and Contributors Information from Bugzilla
by: Acharya, Jagrit, et al.
Published: (2025)
by: Acharya, Jagrit, et al.
Published: (2025)
SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration
by: Chen, Jialong, et al.
Published: (2026)
by: Chen, Jialong, et al.
Published: (2026)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
by: Aleithan, Reem, et al.
Published: (2024)
by: Aleithan, Reem, et al.
Published: (2024)
Challenging Bug Prediction and Repair Models with Synthetic Bugs
by: Ibrahimzada, Ali Reza, et al.
Published: (2023)
by: Ibrahimzada, Ali Reza, et al.
Published: (2023)
Can Old Tests Do New Tricks for Resolving SWE Issues?
by: Chen, Yang, et al.
Published: (2025)
by: Chen, Yang, et al.
Published: (2025)
Beyond Augmentation: Score-Guided Pathological Prior for EEG-based Depression Detection
by: Chen, Xiaojing, et al.
Published: (2026)
by: Chen, Xiaojing, et al.
Published: (2026)
SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents
by: Wang, Yuhang, et al.
Published: (2026)
by: Wang, Yuhang, et al.
Published: (2026)
Accurate background velocity model building method based on iterative deep learning in sparse transform domain
by: Chen, Guoxin
Published: (2024)
by: Chen, Guoxin
Published: (2024)
A Review of Modeling and Waveform Inversion for Marine Seismic Data
by: Chen, Guoxin
Published: (2026)
by: Chen, Guoxin
Published: (2026)
Beyond Single-shot Writing: Deep Research Agents are Unreliable at Multi-turn Report Revision
by: Chen, Bingsen, et al.
Published: (2026)
by: Chen, Bingsen, et al.
Published: (2026)
CalibAnyView: Beyond Single-View Camera Calibration in the Wild
by: Li, Boying, et al.
Published: (2026)
by: Li, Boying, et al.
Published: (2026)
A Comprehensive Study of Bug-Fix Patterns in Autonomous Driving Systems
by: Chen, Yuntianyi, et al.
Published: (2025)
by: Chen, Yuntianyi, et al.
Published: (2025)
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
by: Shen, Chihao, et al.
Published: (2025)
by: Shen, Chihao, et al.
Published: (2025)
ClawGym: A Scalable Framework for Building Effective Claw Agents
by: Bai, Fei, et al.
Published: (2026)
by: Bai, Fei, et al.
Published: (2026)
Fixing Large Language Models' Specification Misunderstanding for Better Code Generation
by: Tian, Zhao, et al.
Published: (2023)
by: Tian, Zhao, et al.
Published: (2023)
Similar Items
-
Toward Autonomous Long-Horizon Engineering for ML Research
by: Chen, Guoxin, et al.
Published: (2026) -
SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training
by: Song, Huatong, et al.
Published: (2026) -
Immersion in the GitHub Universe: Scaling Coding Agents to Mastery
by: Zhao, Jiale, et al.
Published: (2026) -
Computer Environments Elicit General Agentic Intelligence in LLMs
by: Cheng, Daixuan, et al.
Published: (2026) -
SWE-Spot: Building Small Repo-Experts with Repository-Centric Learning
by: Peng, Jinjun, et al.
Published: (2026)