SWE-Bench+: Enhanced Coding Benchmark for LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Aleithan, Reem, Xue, Haoran, Mohajer, Mohammad Mahdi, Nnorom, Elijah, Uddin, Gias, Wang, Song |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
PAGENT: Learning to Patch Software Engineering Agents
par: Xue, Haoran, et autres
Publié: (2025)
par: Xue, Haoran, et autres
Publié: (2025)
LLM Assisted Coding with Metamorphic Specification Mutation Agent
par: Akhond, Mostafijur Rahman, et autres
Publié: (2025)
par: Akhond, Mostafijur Rahman, et autres
Publié: (2025)
StaAgent: An Agentic Framework for Testing Static Analyzers
par: Nnorom, Elijah, et autres
Publié: (2025)
par: Nnorom, Elijah, et autres
Publié: (2025)
Optimized Log Parsing with Syntactic Modifications
par: Enan, Nafid, et autres
Publié: (2025)
par: Enan, Nafid, et autres
Publié: (2025)
Checker Bug Detection and Repair in Deep Learning Libraries
par: Harzevili, Nima Shiri, et autres
Publié: (2024)
par: Harzevili, Nima Shiri, et autres
Publié: (2024)
SWE Context Bench: A Benchmark for Context Learning in Coding
par: Zhu, Jiayuan, et autres
Publié: (2026)
par: Zhu, Jiayuan, et autres
Publié: (2026)
Retrieval-Augmented Test Generation: How Far Are We?
par: Shin, Jiho, et autres
Publié: (2024)
par: Shin, Jiho, et autres
Publié: (2024)
Engineering Pitfalls in AI Coding Tools: An Empirical Study of Bugs in Claude Code, Codex, and Gemini CLI
par: Zhang, Ruixin, et autres
Publié: (2026)
par: Zhang, Ruixin, et autres
Publié: (2026)
ABTest: Behavior-Driven Testing for AI Coding Agents
par: Dai, Wuyang, et autres
Publié: (2026)
par: Dai, Wuyang, et autres
Publié: (2026)
Assessing the Influence of Toxic and Gender Discriminatory Communication on Perceptible Diversity in OSS Projects
par: Sultana, Sayma, et autres
Publié: (2024)
par: Sultana, Sayma, et autres
Publié: (2024)
Evaluating the Environmental Impact of using SLMs and Prompt Engineering for Code Generation
par: Mamun, Md Afif Al, et autres
Publié: (2026)
par: Mamun, Md Afif Al, et autres
Publié: (2026)
An Empirical Study on Bug Severity Estimation using Source Code Metrics and Static Analysis
par: Mashhadi, Ehsan, et autres
Publié: (2022)
par: Mashhadi, Ehsan, et autres
Publié: (2022)
What's in a Benchmark? The Case of SWE-Bench in Automated Program Repair
par: Martinez, Matias, et autres
Publié: (2026)
par: Martinez, Matias, et autres
Publié: (2026)
Bias Testing and Mitigation in Black Box LLMs using Metamorphic Relations
par: Salimian, Sina, et autres
Publié: (2025)
par: Salimian, Sina, et autres
Publié: (2025)
Secret Leak Detection in Software Issue Reports using LLMs: A Comprehensive Evaluation
par: Ahmed, Sadif, et autres
Publié: (2024)
par: Ahmed, Sadif, et autres
Publié: (2024)
SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks
par: Mhatre, Sanket, et autres
Publié: (2025)
par: Mhatre, Sanket, et autres
Publié: (2025)
A Systematic Mapping Study of Crowd Knowledge Enhanced Software Engineering Research Using Stack Overflow
par: Tanzil, Minaoar, et autres
Publié: (2024)
par: Tanzil, Minaoar, et autres
Publié: (2024)
ChatGPT Inaccuracy Mitigation during Technical Report Understanding: Are We There Yet?
par: Tamanna, Salma Begum, et autres
Publié: (2024)
par: Tamanna, Salma Begum, et autres
Publié: (2024)
Reputation Gaming in Stack Overflow
par: Mazloomzadeh, Iren, et autres
Publié: (2021)
par: Mazloomzadeh, Iren, et autres
Publié: (2021)
BLAgent: Agentic RAG for File-Level Bug Localization
par: Mamun, Md Afif Al, et autres
Publié: (2026)
par: Mamun, Md Afif Al, et autres
Publié: (2026)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
par: Jing, Huihao, et autres
Publié: (2026)
par: Jing, Huihao, et autres
Publié: (2026)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
par: Raghavendra, Mohit, et autres
Publié: (2026)
par: Raghavendra, Mohit, et autres
Publié: (2026)
Program Slicing in the Era of Large Language Models
par: Shahandashti, Kimya Khakzad, et autres
Publié: (2024)
par: Shahandashti, Kimya Khakzad, et autres
Publié: (2024)
SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle
par: Guan, Hao, et autres
Publié: (2026)
par: Guan, Hao, et autres
Publié: (2026)
SWE-Bench-CL: Continual Learning for Coding Agents
par: Joshi, Thomas, et autres
Publié: (2025)
par: Joshi, Thomas, et autres
Publié: (2025)
LLM For Loop Invariant Generation and Fixing: How Far Are We?
par: Akhond, Mostafijur Rahman, et autres
Publié: (2025)
par: Akhond, Mostafijur Rahman, et autres
Publié: (2025)
OSS-Bench: Benchmark Generator for Coding LLMs
par: Jiang, Yuancheng, et autres
Publié: (2025)
par: Jiang, Yuancheng, et autres
Publié: (2025)
The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason
par: Liang, Shanchao, et autres
Publié: (2025)
par: Liang, Shanchao, et autres
Publié: (2025)
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
par: Garg, Spandan, et autres
Publié: (2025)
par: Garg, Spandan, et autres
Publié: (2025)
SWE-QA: A Dataset and Benchmark for Complex Code Understanding
par: Elkoussy, Laïla, et autres
Publié: (2026)
par: Elkoussy, Laïla, et autres
Publié: (2026)
From SWE-ZERO to SWE-HERO: Execution-free to Execution-based Fine-tuning for Software Engineering Agents
par: Ludwig, Nikolai, et autres
Publié: (2026)
par: Ludwig, Nikolai, et autres
Publié: (2026)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
par: Xu, Yisen, et autres
Publié: (2026)
par: Xu, Yisen, et autres
Publié: (2026)
Does SWE-Bench-Verified Test Agent Ability or Model Memory?
par: Prathifkumar, Thanosan, et autres
Publié: (2025)
par: Prathifkumar, Thanosan, et autres
Publié: (2025)
"How do people decide?": A Model for Software Library Selection
par: Tanzil, Minaoar Hossain, et autres
Publié: (2024)
par: Tanzil, Minaoar Hossain, et autres
Publié: (2024)
A Large-Scale Empirical Study of COVID-19 Contact Tracing Mobile App Reviews
par: Parisa, Sifat Ishmam, et autres
Publié: (2024)
par: Parisa, Sifat Ishmam, et autres
Publié: (2024)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
par: Chen, Zaoyu, et autres
Publié: (2026)
par: Chen, Zaoyu, et autres
Publié: (2026)
UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench
par: Yu, Boxi, et autres
Publié: (2025)
par: Yu, Boxi, et autres
Publié: (2025)
Evaluating the Effectiveness of GPT-4 Turbo in Creating Defeaters for Assurance Cases
par: Shahandashti, Kimya Khakzad, et autres
Publié: (2024)
par: Shahandashti, Kimya Khakzad, et autres
Publié: (2024)
SWE-Shepherd: Advancing PRMs for Reinforcing Code Agents
par: Dihan, Mahir Labib, et autres
Publié: (2026)
par: Dihan, Mahir Labib, et autres
Publié: (2026)
SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies
par: Saxena, Siddhant, et autres
Publié: (2026)
par: Saxena, Siddhant, et autres
Publié: (2026)
Documents similaires
-
PAGENT: Learning to Patch Software Engineering Agents
par: Xue, Haoran, et autres
Publié: (2025) -
LLM Assisted Coding with Metamorphic Specification Mutation Agent
par: Akhond, Mostafijur Rahman, et autres
Publié: (2025) -
StaAgent: An Agentic Framework for Testing Static Analyzers
par: Nnorom, Elijah, et autres
Publié: (2025) -
Optimized Log Parsing with Syntactic Modifications
par: Enan, Nafid, et autres
Publié: (2025) -
Checker Bug Detection and Repair in Deep Learning Libraries
par: Harzevili, Nima Shiri, et autres
Publié: (2024)