Salvato in:
| Autori principali: | Martinez, Matias, Franch, Xavier |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2602.04449 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Dissecting the SWE-Bench Leaderboards: Profiling Submitters and Architectures of LLM- and Agent-Based Repair Systems
di: Martinez, Matias, et al.
Pubblicazione: (2025)
di: Martinez, Matias, et al.
Pubblicazione: (2025)
Energy Consumption of Automated Program Repair
di: Martinez, Matias, et al.
Pubblicazione: (2022)
di: Martinez, Matias, et al.
Pubblicazione: (2022)
Automated Requirements Relation Extraction
di: Motger, Quim, et al.
Pubblicazione: (2024)
di: Motger, Quim, et al.
Pubblicazione: (2024)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
di: Aleithan, Reem, et al.
Pubblicazione: (2024)
di: Aleithan, Reem, et al.
Pubblicazione: (2024)
SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks
di: Mhatre, Sanket, et al.
Pubblicazione: (2025)
di: Mhatre, Sanket, et al.
Pubblicazione: (2025)
SWE Context Bench: A Benchmark for Context Learning in Coding
di: Zhu, Jiayuan, et al.
Pubblicazione: (2026)
di: Zhu, Jiayuan, et al.
Pubblicazione: (2026)
Cataloguing Hugging Face Models to Software Engineering Activities: Automation and Findings
di: González, Alexandra, et al.
Pubblicazione: (2025)
di: González, Alexandra, et al.
Pubblicazione: (2025)
SEMODS: A Validated Dataset of Open-Source Software Engineering Models
di: González, Alexandra, et al.
Pubblicazione: (2026)
di: González, Alexandra, et al.
Pubblicazione: (2026)
Saving SWE-Bench: A Benchmark Mutation Approach for Realistic Agent Evaluation
di: Garg, Spandan, et al.
Pubblicazione: (2025)
di: Garg, Spandan, et al.
Pubblicazione: (2025)
ThinkRepair: Self-Directed Automated Program Repair
di: Yin, Xin, et al.
Pubblicazione: (2024)
di: Yin, Xin, et al.
Pubblicazione: (2024)
The Impact of Program Reduction on Automated Program Repair
di: Vidziunas, Linas, et al.
Pubblicazione: (2024)
di: Vidziunas, Linas, et al.
Pubblicazione: (2024)
HEJ-Robust: A Robustness Benchmark for LLM-Based Automated Program Repair
di: Rabbi, Fazle, et al.
Pubblicazione: (2026)
di: Rabbi, Fazle, et al.
Pubblicazione: (2026)
ContrastRepair: Enhancing Conversation-Based Automated Program Repair via Contrastive Test Case Pairs
di: Kong, Jiaolong, et al.
Pubblicazione: (2024)
di: Kong, Jiaolong, et al.
Pubblicazione: (2024)
CI-Repair-Bench: A Repository-Aware Benchmark for Automated Patch Validation via CI Workflows
di: Muna, Rabeya Khatun, et al.
Pubblicazione: (2026)
di: Muna, Rabeya Khatun, et al.
Pubblicazione: (2026)
RepairBench: Leaderboard of Frontier Models for Program Repair
di: Silva, André, et al.
Pubblicazione: (2024)
di: Silva, André, et al.
Pubblicazione: (2024)
SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
di: Oliva, Gustavo A., et al.
Pubblicazione: (2025)
di: Oliva, Gustavo A., et al.
Pubblicazione: (2025)
A Tool for Automatically Cataloguing and Selecting Pre-Trained Models and Datasets for Software Engineering
di: González, Alexandra, et al.
Pubblicazione: (2026)
di: González, Alexandra, et al.
Pubblicazione: (2026)
Specification Vibing for Automated Program Repair
di: Zhu, Taohong, et al.
Pubblicazione: (2026)
di: Zhu, Taohong, et al.
Pubblicazione: (2026)
Does SWE-Bench-Verified Test Agent Ability or Model Memory?
di: Prathifkumar, Thanosan, et al.
Pubblicazione: (2025)
di: Prathifkumar, Thanosan, et al.
Pubblicazione: (2025)
Lessons Learned from Mining the Hugging Face Repository
di: Castaño, Joel, et al.
Pubblicazione: (2024)
di: Castaño, Joel, et al.
Pubblicazione: (2024)
Characterizing Datasets for LLM-based Requirements Engineering: A Systematic Mapping Study
di: Motger, Quim, et al.
Pubblicazione: (2025)
di: Motger, Quim, et al.
Pubblicazione: (2025)
Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench
di: Cheshkov, Anton, et al.
Pubblicazione: (2024)
di: Cheshkov, Anton, et al.
Pubblicazione: (2024)
SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
di: Guo, Lianghong, et al.
Pubblicazione: (2025)
di: Guo, Lianghong, et al.
Pubblicazione: (2025)
Innovating for Tomorrow: The Convergence of SE and Green AI
di: Cruz, Luís, et al.
Pubblicazione: (2024)
di: Cruz, Luís, et al.
Pubblicazione: (2024)
What About Emotions? Guiding Fine-Grained Emotion Extraction from Mobile App Reviews
di: Motger, Quim, et al.
Pubblicazione: (2025)
di: Motger, Quim, et al.
Pubblicazione: (2025)
PathFix: Automated Program Repair with Expected Path
di: He, Xu, et al.
Pubblicazione: (2025)
di: He, Xu, et al.
Pubblicazione: (2025)
On The Effectiveness of Dynamic Reduction Techniques in Automated Program Repair
di: Al-Bataineh, Omar I.
Pubblicazione: (2024)
di: Al-Bataineh, Omar I.
Pubblicazione: (2024)
Towards Practical and Useful Automated Program Repair for Debugging
di: Xin, Qi, et al.
Pubblicazione: (2024)
di: Xin, Qi, et al.
Pubblicazione: (2024)
BUGSPHP: A dataset for Automated Program Repair in PHP
di: Pramod, K. D., et al.
Pubblicazione: (2024)
di: Pramod, K. D., et al.
Pubblicazione: (2024)
ASAP-Repair: API-Specific Automated Program Repair Based on API Usage Graphs
di: Nielebock, Sebastian, et al.
Pubblicazione: (2024)
di: Nielebock, Sebastian, et al.
Pubblicazione: (2024)
Unveiling Competition Dynamics in Mobile App Markets through User Reviews
di: Motger, Quim, et al.
Pubblicazione: (2023)
di: Motger, Quim, et al.
Pubblicazione: (2023)
Multi-Agent Debate Strategies to Enhance Requirements Engineering with Large Language Models
di: Oriol, Marc, et al.
Pubblicazione: (2025)
di: Oriol, Marc, et al.
Pubblicazione: (2025)
SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents
di: Rashid, Muhammad Shihab, et al.
Pubblicazione: (2025)
di: Rashid, Muhammad Shihab, et al.
Pubblicazione: (2025)
How Safe Are AI-Generated Patches? A Large-scale Study on Security Risks in LLM and Agentic Automated Program Repair on SWE-bench
di: Sajadi, Amirali, et al.
Pubblicazione: (2025)
di: Sajadi, Amirali, et al.
Pubblicazione: (2025)
Hybrid Automated Program Repair by Combining Large Language Models and Program Analysis
di: Li, Fengjie, et al.
Pubblicazione: (2024)
di: Li, Fengjie, et al.
Pubblicazione: (2024)
Software-Based Dialogue Systems: Survey, Taxonomy and Challenges
di: Motger, Quim, et al.
Pubblicazione: (2021)
di: Motger, Quim, et al.
Pubblicazione: (2021)
A Methodological Framework for LLM-Based Mining of Software Repositories
di: De Martino, Vincenzo, et al.
Pubblicazione: (2025)
di: De Martino, Vincenzo, et al.
Pubblicazione: (2025)
A Framework for Using LLMs for Repository Mining Studies in Empirical Software Engineering
di: de Martino, Vincenzo, et al.
Pubblicazione: (2024)
di: de Martino, Vincenzo, et al.
Pubblicazione: (2024)
Automated Test Case Repair Using Language Models
di: Yaraghi, Ahmadreza Saboor, et al.
Pubblicazione: (2024)
di: Yaraghi, Ahmadreza Saboor, et al.
Pubblicazione: (2024)
Automated Repair of C Programs Using Large Language Models
di: Farzandway, Mahdi, et al.
Pubblicazione: (2025)
di: Farzandway, Mahdi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Dissecting the SWE-Bench Leaderboards: Profiling Submitters and Architectures of LLM- and Agent-Based Repair Systems
di: Martinez, Matias, et al.
Pubblicazione: (2025) -
Energy Consumption of Automated Program Repair
di: Martinez, Matias, et al.
Pubblicazione: (2022) -
Automated Requirements Relation Extraction
di: Motger, Quim, et al.
Pubblicazione: (2024) -
SWE-Bench+: Enhanced Coding Benchmark for LLMs
di: Aleithan, Reem, et al.
Pubblicazione: (2024) -
SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks
di: Mhatre, Sanket, et al.
Pubblicazione: (2025)