SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mündler, Niels, Müller, Mark Niklas, He, Jingxuan, Vechev, Martin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
von: Thillen, Alex, et al.
Veröffentlicht: (2026)
von: Thillen, Alex, et al.
Veröffentlicht: (2026)
Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
Black-Box Adversarial Attacks on LLM-Based Code Completion
von: Jenko, Slobodan, et al.
Veröffentlicht: (2024)
von: Jenko, Slobodan, et al.
Veröffentlicht: (2024)
Instruction Tuning for Secure Code Generation
von: He, Jingxuan, et al.
Veröffentlicht: (2024)
von: He, Jingxuan, et al.
Veröffentlicht: (2024)
Coding Agents Don't Know When to Act
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
Large Language Models for Code: Security Hardening and Adversarial Testing
von: He, Jingxuan, et al.
Veröffentlicht: (2023)
von: He, Jingxuan, et al.
Veröffentlicht: (2023)
Automated Benchmark Generation for Repository-Level Coding Tasks
von: Vergopoulos, Konstantinos, et al.
Veröffentlicht: (2025)
von: Vergopoulos, Konstantinos, et al.
Veröffentlicht: (2025)
On the Impact of Code Comments for Automated Bug-Fixing: An Empirical Study
von: Vitale, Antonio, et al.
Veröffentlicht: (2026)
von: Vitale, Antonio, et al.
Veröffentlicht: (2026)
Constrained Decoding of Diffusion LLMs with Context-Free Grammars
von: Mündler, Niels, et al.
Veröffentlicht: (2025)
von: Mündler, Niels, et al.
Veröffentlicht: (2025)
FlakyFix: Using Large Language Models for Predicting Flaky Test Fix Categories and Test Code Repair
von: Fatima, Sakina, et al.
Veröffentlicht: (2023)
von: Fatima, Sakina, et al.
Veröffentlicht: (2023)
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
von: Mündler, Niels, et al.
Veröffentlicht: (2023)
von: Mündler, Niels, et al.
Veröffentlicht: (2023)
SWE-Bench-CL: Continual Learning for Coding Agents
von: Joshi, Thomas, et al.
Veröffentlicht: (2025)
von: Joshi, Thomas, et al.
Veröffentlicht: (2025)
PerfBench: Can Agents Resolve Real-World Performance Bugs?
von: Garg, Spandan, et al.
Veröffentlicht: (2025)
von: Garg, Spandan, et al.
Veröffentlicht: (2025)
MarsCode Agent: AI-native Automated Bug Fixing
von: Liu, Yizhou, et al.
Veröffentlicht: (2024)
von: Liu, Yizhou, et al.
Veröffentlicht: (2024)
Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases
von: Wong, Sherman, et al.
Veröffentlicht: (2025)
von: Wong, Sherman, et al.
Veröffentlicht: (2025)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
von: Vulićević, Jelena Ilić
Veröffentlicht: (2026)
von: Vulićević, Jelena Ilić
Veröffentlicht: (2026)
Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent Agent
von: Shah, Mehil B, et al.
Veröffentlicht: (2025)
von: Shah, Mehil B, et al.
Veröffentlicht: (2025)
Debug2Fix: Can Interactive Debugging Help Coding Agents Fix More Bugs?
von: Garg, Spandan, et al.
Veröffentlicht: (2026)
von: Garg, Spandan, et al.
Veröffentlicht: (2026)
scicode-lint: Detecting Methodology Bugs in Scientific Python Code with LLM-Generated Patterns
von: Samsonau, Sergey V.
Veröffentlicht: (2026)
von: Samsonau, Sergey V.
Veröffentlicht: (2026)
Code Generation by Differential Test Time Scaling
von: He, Yifeng, et al.
Veröffentlicht: (2026)
von: He, Yifeng, et al.
Veröffentlicht: (2026)
SWE-Synth: Synthesizing Verifiable Bug-Fix Data to Enable Large Language Models in Resolving Real-World Bugs
von: Pham, Minh V. T., et al.
Veröffentlicht: (2025)
von: Pham, Minh V. T., et al.
Veröffentlicht: (2025)
MobiFlow: Real-World Mobile Agent Benchmarking through Trajectory Fusion
von: Feng, Yunfei, et al.
Veröffentlicht: (2026)
von: Feng, Yunfei, et al.
Veröffentlicht: (2026)
ToolFuzz -- Automated Agent Tool Testing
von: Milev, Ivan, et al.
Veröffentlicht: (2025)
von: Milev, Ivan, et al.
Veröffentlicht: (2025)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
von: Rahman, Musfiqur, et al.
Veröffentlicht: (2025)
von: Rahman, Musfiqur, et al.
Veröffentlicht: (2025)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
von: Daghighfarsoodeh, Alireza, et al.
Veröffentlicht: (2025)
von: Daghighfarsoodeh, Alireza, et al.
Veröffentlicht: (2025)
First Three Years of the International Verification of Neural Networks Competition (VNN-COMP)
von: Brix, Christopher, et al.
Veröffentlicht: (2023)
von: Brix, Christopher, et al.
Veröffentlicht: (2023)
Can Coding Agents Be General Agents?
von: Ivanov, Maksim, et al.
Veröffentlicht: (2026)
von: Ivanov, Maksim, et al.
Veröffentlicht: (2026)
Are Large Language Models Memorizing Bug Benchmarks?
von: Ramos, Daniel, et al.
Veröffentlicht: (2024)
von: Ramos, Daniel, et al.
Veröffentlicht: (2024)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
Are Sparse Autoencoders Useful for Java Function Bug Detection?
von: Melo, Rui, et al.
Veröffentlicht: (2025)
von: Melo, Rui, et al.
Veröffentlicht: (2025)
Co-Located Tests, Better AI Code: How Test Syntax Structure Affects Foundation Model Code Generation
von: Jacopin, Éric
Veröffentlicht: (2026)
von: Jacopin, Éric
Veröffentlicht: (2026)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
SetupBench: Assessing Software Engineering Agents' Ability to Bootstrap Development Environments
von: Arora, Avi, et al.
Veröffentlicht: (2025)
von: Arora, Avi, et al.
Veröffentlicht: (2025)
PostTrainBench: Can LLM Agents Automate LLM Post-Training?
von: Rank, Ben, et al.
Veröffentlicht: (2026)
von: Rank, Ben, et al.
Veröffentlicht: (2026)
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
von: Xiao, Yijia, et al.
Veröffentlicht: (2025)
von: Xiao, Yijia, et al.
Veröffentlicht: (2025)
Deploying Geospatial Foundation Models in the Real World: Lessons from WorldCereal
von: Butsko, Christina, et al.
Veröffentlicht: (2025)
von: Butsko, Christina, et al.
Veröffentlicht: (2025)
An Empirical Study on LLM-based Agents for Automated Bug Fixing
von: Meng, Xiangxin, et al.
Veröffentlicht: (2024)
von: Meng, Xiangxin, et al.
Veröffentlicht: (2024)
A Theoretical Analysis of Test-Driven Code Generation
von: Menet, Nicolas, et al.
Veröffentlicht: (2026)
von: Menet, Nicolas, et al.
Veröffentlicht: (2026)
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
von: Trivedi, Harsh, et al.
Veröffentlicht: (2024)
von: Trivedi, Harsh, et al.
Veröffentlicht: (2024)
GREPO: A Benchmark for Graph Neural Networks on Repository-Level Bug Localization
von: Wang, Juntong, et al.
Veröffentlicht: (2026)
von: Wang, Juntong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
von: Thillen, Alex, et al.
Veröffentlicht: (2026) -
Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026) -
Black-Box Adversarial Attacks on LLM-Based Code Completion
von: Jenko, Slobodan, et al.
Veröffentlicht: (2024) -
Instruction Tuning for Secure Code Generation
von: He, Jingxuan, et al.
Veröffentlicht: (2024) -
Coding Agents Don't Know When to Act
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)