SWE-Bench++: A Framework for the Scalable Generation of Software Engineering Benchmarks from Open-Source Repositories
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Lilin, Ramalho, Lucas, Celestino, Alan, Pham, Phuc Anthony, Liu, Yu, Sinha, Umang Kumar, Portillo, Andres, Osunwa, Onassis, Maduekwe, Gabriel |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks
par: Mhatre, Sanket, et autres
Publié: (2025)
par: Mhatre, Sanket, et autres
Publié: (2025)
AbasyBench
par: Pham, Bao Phuc, et autres
Publié: (2026)
par: Pham, Bao Phuc, et autres
Publié: (2026)
SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories
par: Soni, Aditya Bharat, et autres
Publié: (2026)
par: Soni, Aditya Bharat, et autres
Publié: (2026)
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
par: Liang, Jiarong, et autres
Publié: (2026)
par: Liang, Jiarong, et autres
Publié: (2026)
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
par: Cai, Songcheng, et autres
Publié: (2026)
par: Cai, Songcheng, et autres
Publié: (2026)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
par: Aleithan, Reem, et autres
Publié: (2024)
par: Aleithan, Reem, et autres
Publié: (2024)
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
par: Deng, Xiang, et autres
Publié: (2025)
par: Deng, Xiang, et autres
Publié: (2025)
SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?
par: Han, Tingxu, et autres
Publié: (2026)
par: Han, Tingxu, et autres
Publié: (2026)
SWE-Hub: A Unified Production System for Scalable, Executable Software Engineering Tasks
par: Zeng, Yucheng, et autres
Publié: (2026)
par: Zeng, Yucheng, et autres
Publié: (2026)
SWE-Bench 5G: Benchmarking AI Coding Agents on Telecom Network Engineering Tasks
par: Chen, Jiao, et autres
Publié: (2026)
par: Chen, Jiao, et autres
Publié: (2026)
ATime-Consistent Benchmark for Repository-Level Software Engineering Evaluation
par: Xianpeng, et autres
Publié: (2026)
par: Xianpeng, et autres
Publié: (2026)
SWE-smith: Scaling Data for Software Engineering Agents
par: Yang, John, et autres
Publié: (2025)
par: Yang, John, et autres
Publié: (2025)
Training Software Engineering Agents and Verifiers with SWE-Gym
par: Pan, Jiayi, et autres
Publié: (2024)
par: Pan, Jiayi, et autres
Publié: (2024)
SWE Context Bench: A Benchmark for Context Learning in Coding
par: Zhu, Jiayuan, et autres
Publié: (2026)
par: Zhu, Jiayuan, et autres
Publié: (2026)
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
par: Zhang, Zehua, et autres
Publié: (2025)
par: Zhang, Zehua, et autres
Publié: (2025)
Revealing the value of Repository Centrality in lifespan prediction of Open Source Software Projects
par: He, Runzhi, et autres
Publié: (2024)
par: He, Runzhi, et autres
Publié: (2024)
TOM-SWE: User Mental Modeling For Software Engineering Agents
par: Zhou, Xuhui, et autres
Publié: (2025)
par: Zhou, Xuhui, et autres
Publié: (2025)
SWE-RM: Execution-free Feedback For Software Engineering Agents
par: Shum, KaShun, et autres
Publié: (2025)
par: Shum, KaShun, et autres
Publié: (2025)
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
par: Vijayvargiya, Sanidhya, et autres
Publié: (2025)
par: Vijayvargiya, Sanidhya, et autres
Publié: (2025)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
par: Xu, Yisen, et autres
Publié: (2026)
par: Xu, Yisen, et autres
Publié: (2026)
What's in a Benchmark? The Case of SWE-Bench in Automated Program Repair
par: Martinez, Matias, et autres
Publié: (2026)
par: Martinez, Matias, et autres
Publié: (2026)
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
par: Adamenko, Pavel, et autres
Publié: (2025)
par: Adamenko, Pavel, et autres
Publié: (2025)
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
par: Wei, Yuxiang, et autres
Publié: (2025)
par: Wei, Yuxiang, et autres
Publié: (2025)
From SWE-ZERO to SWE-HERO: Execution-free to Execution-based Fine-tuning for Software Engineering Agents
par: Ludwig, Nikolai, et autres
Publié: (2026)
par: Ludwig, Nikolai, et autres
Publié: (2026)
Exploring the Lifecycle and Maintenance Practices of Pre-Trained Models in Open-Source Software Repositories
par: Koohjani, Matin, et autres
Publié: (2025)
par: Koohjani, Matin, et autres
Publié: (2025)
Mining Quantum Software Patterns in Open-Source Projects
par: Ramalho, Neilson Carlos Leite, et autres
Publié: (2026)
par: Ramalho, Neilson Carlos Leite, et autres
Publié: (2026)
Resolving Java Code Repository Issues with iSWE Agent
par: Ganhotra, Jatin, et autres
Publié: (2026)
par: Ganhotra, Jatin, et autres
Publié: (2026)
NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
par: Shao, Minghao, et autres
Publié: (2024)
par: Shao, Minghao, et autres
Publié: (2024)
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
par: Garg, Spandan, et autres
Publié: (2025)
par: Garg, Spandan, et autres
Publié: (2025)
SWE-Arena: An Interactive Platform for Evaluating Foundation Models in Software Engineering
par: Zhao, Zhimin
Publié: (2025)
par: Zhao, Zhimin
Publié: (2025)
SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling
par: Wang, Haoran, et autres
Publié: (2025)
par: Wang, Haoran, et autres
Publié: (2025)
Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs
par: Zeng, Liang, et autres
Publié: (2025)
par: Zeng, Liang, et autres
Publié: (2025)
SWE-World: Building Software Engineering Agents in Docker-Free Environments
par: Sun, Shuang, et autres
Publié: (2026)
par: Sun, Shuang, et autres
Publié: (2026)
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
par: Yang, John, et autres
Publié: (2024)
par: Yang, John, et autres
Publié: (2024)
SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents
par: Ding, Yifeng, et autres
Publié: (2026)
par: Ding, Yifeng, et autres
Publié: (2026)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
par: Duston, Titouan, et autres
Publié: (2025)
par: Duston, Titouan, et autres
Publié: (2025)
SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies
par: Saxena, Siddhant, et autres
Publié: (2026)
par: Saxena, Siddhant, et autres
Publié: (2026)
SWE-Spot: Building Small Repo-Experts with Repository-Centric Learning
par: Peng, Jinjun, et autres
Publié: (2026)
par: Peng, Jinjun, et autres
Publié: (2026)
SWE-Flow: Synthesizing Software Engineering Data in a Test-Driven Manner
par: Zhang, Lei, et autres
Publié: (2025)
par: Zhang, Lei, et autres
Publié: (2025)
Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering
par: Zeng, Guangtao, et autres
Publié: (2025)
par: Zeng, Guangtao, et autres
Publié: (2025)
Documents similaires
-
SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks
par: Mhatre, Sanket, et autres
Publié: (2025) -
AbasyBench
par: Pham, Bao Phuc, et autres
Publié: (2026) -
SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories
par: Soni, Aditya Bharat, et autres
Publié: (2026) -
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
par: Liang, Jiarong, et autres
Publié: (2026) -
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
par: Cai, Songcheng, et autres
Publié: (2026)