From SWE-ZERO to SWE-HERO: Execution-free to Execution-based Fine-tuning for Software Engineering Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ludwig, Nikolai, Ahmad, Wasi Uddin, Majumdar, Somshubra, Ginsburg, Boris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Output to Evaluation: Does Raw Instruction-Tuned Code LLMs Output Suffice for Fill-in-the-Middle Code Generation?
von: Ahmad, Wasi Uddin, et al.
Veröffentlicht: (2025)
von: Ahmad, Wasi Uddin, et al.
Veröffentlicht: (2025)
OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs
von: Ahmad, Wasi Uddin, et al.
Veröffentlicht: (2025)
von: Ahmad, Wasi Uddin, et al.
Veröffentlicht: (2025)
SWE-smith: Scaling Data for Software Engineering Agents
von: Yang, John, et al.
Veröffentlicht: (2025)
von: Yang, John, et al.
Veröffentlicht: (2025)
Training Software Engineering Agents and Verifiers with SWE-Gym
von: Pan, Jiayi, et al.
Veröffentlicht: (2024)
von: Pan, Jiayi, et al.
Veröffentlicht: (2024)
SWE-Lego: Pushing the Limits of Supervised Fine-tuning for Software Issue Resolving
von: Tao, Chaofan, et al.
Veröffentlicht: (2026)
von: Tao, Chaofan, et al.
Veröffentlicht: (2026)
SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
von: Badertdinov, Ibragim, et al.
Veröffentlicht: (2025)
von: Badertdinov, Ibragim, et al.
Veröffentlicht: (2025)
SWE-World: Building Software Engineering Agents in Docker-Free Environments
von: Sun, Shuang, et al.
Veröffentlicht: (2026)
von: Sun, Shuang, et al.
Veröffentlicht: (2026)
SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training
von: Song, Huatong, et al.
Veröffentlicht: (2026)
von: Song, Huatong, et al.
Veröffentlicht: (2026)
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
von: Deng, Xiang, et al.
Veröffentlicht: (2025)
von: Deng, Xiang, et al.
Veröffentlicht: (2025)
SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale
von: Badertdinov, Ibragim, et al.
Veröffentlicht: (2026)
von: Badertdinov, Ibragim, et al.
Veröffentlicht: (2026)
SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution
von: Li, Han, et al.
Veröffentlicht: (2025)
von: Li, Han, et al.
Veröffentlicht: (2025)
APEX-SWE
von: Kottamasu, Abhi, et al.
Veröffentlicht: (2026)
von: Kottamasu, Abhi, et al.
Veröffentlicht: (2026)
Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2025)
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2025)
Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning
von: Ficek, Aleksander, et al.
Veröffentlicht: (2025)
von: Ficek, Aleksander, et al.
Veröffentlicht: (2025)
SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development
von: Du, Yaxin, et al.
Veröffentlicht: (2025)
von: Du, Yaxin, et al.
Veröffentlicht: (2025)
Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering
von: Zeng, Guangtao, et al.
Veröffentlicht: (2025)
von: Zeng, Guangtao, et al.
Veröffentlicht: (2025)
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
von: Yang, John, et al.
Veröffentlicht: (2024)
von: Yang, John, et al.
Veröffentlicht: (2024)
SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents
von: Wang, Yuhang, et al.
Veröffentlicht: (2026)
von: Wang, Yuhang, et al.
Veröffentlicht: (2026)
SWE-Exp: Experience-Driven Software Issue Resolution
von: Chen, Silin, et al.
Veröffentlicht: (2025)
von: Chen, Silin, et al.
Veröffentlicht: (2025)
GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents
von: Shetty, Manish, et al.
Veröffentlicht: (2025)
von: Shetty, Manish, et al.
Veröffentlicht: (2025)
SWE-Hub: A Unified Production System for Scalable, Executable Software Engineering Tasks
von: Zeng, Yucheng, et al.
Veröffentlicht: (2026)
von: Zeng, Yucheng, et al.
Veröffentlicht: (2026)
SWE-RM: Execution-free Feedback For Software Engineering Agents
von: Shum, KaShun, et al.
Veröffentlicht: (2025)
von: Shum, KaShun, et al.
Veröffentlicht: (2025)
SWE-bench Goes Live!
von: Zhang, Linghao, et al.
Veröffentlicht: (2025)
von: Zhang, Linghao, et al.
Veröffentlicht: (2025)
SWE-AGI: Benchmarking Specification-Driven Software Construction with MoonBit in the Era of Autonomous Agents
von: Zhang, Zhirui, et al.
Veröffentlicht: (2026)
von: Zhang, Zhirui, et al.
Veröffentlicht: (2026)
SWE-Protégé: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents
von: Kon, Patrick Tser Jern, et al.
Veröffentlicht: (2026)
von: Kon, Patrick Tser Jern, et al.
Veröffentlicht: (2026)
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
von: Adamenko, Pavel, et al.
Veröffentlicht: (2025)
von: Adamenko, Pavel, et al.
Veröffentlicht: (2025)
Toward Training Superintelligent Software Agents through Self-Play SWE-RL
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents
von: Yang, Zonghan, et al.
Veröffentlicht: (2025)
von: Yang, Zonghan, et al.
Veröffentlicht: (2025)
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
von: Yang, John, et al.
Veröffentlicht: (2024)
von: Yang, John, et al.
Veröffentlicht: (2024)
BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?
von: Chen, Guoxin, et al.
Veröffentlicht: (2026)
von: Chen, Guoxin, et al.
Veröffentlicht: (2026)
ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents
von: Li, Kenan, et al.
Veröffentlicht: (2026)
von: Li, Kenan, et al.
Veröffentlicht: (2026)
Heterogeneous Prompting and Execution Feedback for SWE Issue Test Generation and Selection
von: Ahmed, Toufique, et al.
Veröffentlicht: (2025)
von: Ahmed, Toufique, et al.
Veröffentlicht: (2025)
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench
von: Yu, Boxi, et al.
Veröffentlicht: (2025)
von: Yu, Boxi, et al.
Veröffentlicht: (2025)
SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration
von: Chen, Jialong, et al.
Veröffentlicht: (2026)
von: Chen, Jialong, et al.
Veröffentlicht: (2026)
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
von: Lam, Man Ho, et al.
Veröffentlicht: (2026)
von: Lam, Man Ho, et al.
Veröffentlicht: (2026)
SWE-Bench++: A Framework for the Scalable Generation of Software Engineering Benchmarks from Open-Source Repositories
von: Wang, Lilin, et al.
Veröffentlicht: (2025)
von: Wang, Lilin, et al.
Veröffentlicht: (2025)
TOM-SWE: User Mental Modeling For Software Engineering Agents
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2025)
SWE-QA: Can Language Models Answer Repository-level Code Questions?
von: Peng, Weihan, et al.
Veröffentlicht: (2025)
von: Peng, Weihan, et al.
Veröffentlicht: (2025)
Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
von: Zan, Daoguang, et al.
Veröffentlicht: (2025)
von: Zan, Daoguang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
From Output to Evaluation: Does Raw Instruction-Tuned Code LLMs Output Suffice for Fill-in-the-Middle Code Generation?
von: Ahmad, Wasi Uddin, et al.
Veröffentlicht: (2025) -
OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs
von: Ahmad, Wasi Uddin, et al.
Veröffentlicht: (2025) -
SWE-smith: Scaling Data for Software Engineering Agents
von: Yang, John, et al.
Veröffentlicht: (2025) -
Training Software Engineering Agents and Verifiers with SWE-Gym
von: Pan, Jiayi, et al.
Veröffentlicht: (2024) -
SWE-Lego: Pushing the Limits of Supervised Fine-tuning for Software Issue Resolving
von: Tao, Chaofan, et al.
Veröffentlicht: (2026)