Gespeichert in:
| Hauptverfasser: | Martinez, Matias, Franch, Xavier |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.04449 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dissecting the SWE-Bench Leaderboards: Profiling Submitters and Architectures of LLM- and Agent-Based Repair Systems
von: Martinez, Matias, et al.
Veröffentlicht: (2025)
von: Martinez, Matias, et al.
Veröffentlicht: (2025)
Energy Consumption of Automated Program Repair
von: Martinez, Matias, et al.
Veröffentlicht: (2022)
von: Martinez, Matias, et al.
Veröffentlicht: (2022)
Automated Requirements Relation Extraction
von: Motger, Quim, et al.
Veröffentlicht: (2024)
von: Motger, Quim, et al.
Veröffentlicht: (2024)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
von: Aleithan, Reem, et al.
Veröffentlicht: (2024)
SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks
von: Mhatre, Sanket, et al.
Veröffentlicht: (2025)
von: Mhatre, Sanket, et al.
Veröffentlicht: (2025)
SWE Context Bench: A Benchmark for Context Learning in Coding
von: Zhu, Jiayuan, et al.
Veröffentlicht: (2026)
von: Zhu, Jiayuan, et al.
Veröffentlicht: (2026)
Cataloguing Hugging Face Models to Software Engineering Activities: Automation and Findings
von: González, Alexandra, et al.
Veröffentlicht: (2025)
von: González, Alexandra, et al.
Veröffentlicht: (2025)
SEMODS: A Validated Dataset of Open-Source Software Engineering Models
von: González, Alexandra, et al.
Veröffentlicht: (2026)
von: González, Alexandra, et al.
Veröffentlicht: (2026)
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
von: Garg, Spandan, et al.
Veröffentlicht: (2025)
von: Garg, Spandan, et al.
Veröffentlicht: (2025)
ThinkRepair: Self-Directed Automated Program Repair
von: Yin, Xin, et al.
Veröffentlicht: (2024)
von: Yin, Xin, et al.
Veröffentlicht: (2024)
The Impact of Program Reduction on Automated Program Repair
von: Vidziunas, Linas, et al.
Veröffentlicht: (2024)
von: Vidziunas, Linas, et al.
Veröffentlicht: (2024)
HEJ-Robust: A Robustness Benchmark for LLM-Based Automated Program Repair
von: Rabbi, Fazle, et al.
Veröffentlicht: (2026)
von: Rabbi, Fazle, et al.
Veröffentlicht: (2026)
ContrastRepair: Enhancing Conversation-Based Automated Program Repair via Contrastive Test Case Pairs
von: Kong, Jiaolong, et al.
Veröffentlicht: (2024)
von: Kong, Jiaolong, et al.
Veröffentlicht: (2024)
CI-Repair-Bench: A Repository-Aware Benchmark for Automated Patch Validation via CI Workflows
von: Muna, Rabeya Khatun, et al.
Veröffentlicht: (2026)
von: Muna, Rabeya Khatun, et al.
Veröffentlicht: (2026)
RepairBench: Leaderboard of Frontier Models for Program Repair
von: Silva, André, et al.
Veröffentlicht: (2024)
von: Silva, André, et al.
Veröffentlicht: (2024)
SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
von: Oliva, Gustavo A., et al.
Veröffentlicht: (2025)
von: Oliva, Gustavo A., et al.
Veröffentlicht: (2025)
A Tool for Automatically Cataloguing and Selecting Pre-Trained Models and Datasets for Software Engineering
von: González, Alexandra, et al.
Veröffentlicht: (2026)
von: González, Alexandra, et al.
Veröffentlicht: (2026)
Specification Vibing for Automated Program Repair
von: Zhu, Taohong, et al.
Veröffentlicht: (2026)
von: Zhu, Taohong, et al.
Veröffentlicht: (2026)
Does SWE-Bench-Verified Test Agent Ability or Model Memory?
von: Prathifkumar, Thanosan, et al.
Veröffentlicht: (2025)
von: Prathifkumar, Thanosan, et al.
Veröffentlicht: (2025)
Lessons Learned from Mining the Hugging Face Repository
von: Castaño, Joel, et al.
Veröffentlicht: (2024)
von: Castaño, Joel, et al.
Veröffentlicht: (2024)
Characterizing Datasets for LLM-based Requirements Engineering: A Systematic Mapping Study
von: Motger, Quim, et al.
Veröffentlicht: (2025)
von: Motger, Quim, et al.
Veröffentlicht: (2025)
Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench
von: Cheshkov, Anton, et al.
Veröffentlicht: (2024)
von: Cheshkov, Anton, et al.
Veröffentlicht: (2024)
SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
von: Guo, Lianghong, et al.
Veröffentlicht: (2025)
von: Guo, Lianghong, et al.
Veröffentlicht: (2025)
Innovating for Tomorrow: The Convergence of SE and Green AI
von: Cruz, Luís, et al.
Veröffentlicht: (2024)
von: Cruz, Luís, et al.
Veröffentlicht: (2024)
What About Emotions? Guiding Fine-Grained Emotion Extraction from Mobile App Reviews
von: Motger, Quim, et al.
Veröffentlicht: (2025)
von: Motger, Quim, et al.
Veröffentlicht: (2025)
PathFix: Automated Program Repair with Expected Path
von: He, Xu, et al.
Veröffentlicht: (2025)
von: He, Xu, et al.
Veröffentlicht: (2025)
On The Effectiveness of Dynamic Reduction Techniques in Automated Program Repair
von: Al-Bataineh, Omar I.
Veröffentlicht: (2024)
von: Al-Bataineh, Omar I.
Veröffentlicht: (2024)
Towards Practical and Useful Automated Program Repair for Debugging
von: Xin, Qi, et al.
Veröffentlicht: (2024)
von: Xin, Qi, et al.
Veröffentlicht: (2024)
BUGSPHP: A dataset for Automated Program Repair in PHP
von: Pramod, K. D., et al.
Veröffentlicht: (2024)
von: Pramod, K. D., et al.
Veröffentlicht: (2024)
ASAP-Repair: API-Specific Automated Program Repair Based on API Usage Graphs
von: Nielebock, Sebastian, et al.
Veröffentlicht: (2024)
von: Nielebock, Sebastian, et al.
Veröffentlicht: (2024)
Unveiling Competition Dynamics in Mobile App Markets through User Reviews
von: Motger, Quim, et al.
Veröffentlicht: (2023)
von: Motger, Quim, et al.
Veröffentlicht: (2023)
Multi-Agent Debate Strategies to Enhance Requirements Engineering with Large Language Models
von: Oriol, Marc, et al.
Veröffentlicht: (2025)
von: Oriol, Marc, et al.
Veröffentlicht: (2025)
SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents
von: Rashid, Muhammad Shihab, et al.
Veröffentlicht: (2025)
von: Rashid, Muhammad Shihab, et al.
Veröffentlicht: (2025)
How Safe Are AI-Generated Patches? A Large-scale Study on Security Risks in LLM and Agentic Automated Program Repair on SWE-bench
von: Sajadi, Amirali, et al.
Veröffentlicht: (2025)
von: Sajadi, Amirali, et al.
Veröffentlicht: (2025)
Hybrid Automated Program Repair by Combining Large Language Models and Program Analysis
von: Li, Fengjie, et al.
Veröffentlicht: (2024)
von: Li, Fengjie, et al.
Veröffentlicht: (2024)
Software-Based Dialogue Systems: Survey, Taxonomy and Challenges
von: Motger, Quim, et al.
Veröffentlicht: (2021)
von: Motger, Quim, et al.
Veröffentlicht: (2021)
A Methodological Framework for LLM-Based Mining of Software Repositories
von: De Martino, Vincenzo, et al.
Veröffentlicht: (2025)
von: De Martino, Vincenzo, et al.
Veröffentlicht: (2025)
A Framework for Using LLMs for Repository Mining Studies in Empirical Software Engineering
von: de Martino, Vincenzo, et al.
Veröffentlicht: (2024)
von: de Martino, Vincenzo, et al.
Veröffentlicht: (2024)
Automated Test Case Repair Using Language Models
von: Yaraghi, Ahmadreza Saboor, et al.
Veröffentlicht: (2024)
von: Yaraghi, Ahmadreza Saboor, et al.
Veröffentlicht: (2024)
Automated Repair of C Programs Using Large Language Models
von: Farzandway, Mahdi, et al.
Veröffentlicht: (2025)
von: Farzandway, Mahdi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Dissecting the SWE-Bench Leaderboards: Profiling Submitters and Architectures of LLM- and Agent-Based Repair Systems
von: Martinez, Matias, et al.
Veröffentlicht: (2025) -
Energy Consumption of Automated Program Repair
von: Martinez, Matias, et al.
Veröffentlicht: (2022) -
Automated Requirements Relation Extraction
von: Motger, Quim, et al.
Veröffentlicht: (2024) -
SWE-Bench+: Enhanced Coding Benchmark for LLMs
von: Aleithan, Reem, et al.
Veröffentlicht: (2024) -
SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks
von: Mhatre, Sanket, et al.
Veröffentlicht: (2025)