From Bugs to Benchmarks: A Comprehensive Survey of Software Defect Datasets
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Hao-Nan, Furth, Robert M., Pradel, Michael, Rubio-González, Cindy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DyPyBench: A Benchmark of Executable Python Software
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
AgentStepper: Interactive Debugging of Software Development Agents
von: Hutter, Robert, et al.
Veröffentlicht: (2026)
von: Hutter, Robert, et al.
Veröffentlicht: (2026)
Evaluating LLM Agents on Automated Software Analysis Tasks
von: Bouzenia, Islem, et al.
Veröffentlicht: (2026)
von: Bouzenia, Islem, et al.
Veröffentlicht: (2026)
Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories
von: Bouzenia, Islem, et al.
Veröffentlicht: (2025)
von: Bouzenia, Islem, et al.
Veröffentlicht: (2025)
BioDefect: The First Dataset for Defect Detection in Bioinformatics Software
von: Xu, Tianxiang, et al.
Veröffentlicht: (2026)
von: Xu, Tianxiang, et al.
Veröffentlicht: (2026)
Testora: Using Natural Language Intent to Detect Behavioral Regressions
von: Pradel, Michael
Veröffentlicht: (2025)
von: Pradel, Michael
Veröffentlicht: (2025)
BugsRepo: A Comprehensive Curated Dataset of Bug Reports, Comments and Contributors Information from Bugzilla
von: Acharya, Jagrit, et al.
Veröffentlicht: (2025)
von: Acharya, Jagrit, et al.
Veröffentlicht: (2025)
A Comprehensive Survey of Benchmarks for Automated Improvement of Software's Non-Functional Properties
von: Blot, Aymeric, et al.
Veröffentlicht: (2022)
von: Blot, Aymeric, et al.
Veröffentlicht: (2022)
A Survey on Testing and Analysis of Quantum Software
von: Paltenghi, Matteo, et al.
Veröffentlicht: (2024)
von: Paltenghi, Matteo, et al.
Veröffentlicht: (2024)
CodeMapper: A Language-Agnostic Approach to Mapping Code Regions Across Commits
von: Hu, Huimin, et al.
Veröffentlicht: (2025)
von: Hu, Huimin, et al.
Veröffentlicht: (2025)
An Empirical Study on Embodied Artificial Intelligence Robot (EAIR) Software Bugs
von: Liao, Zeqin, et al.
Veröffentlicht: (2025)
von: Liao, Zeqin, et al.
Veröffentlicht: (2025)
De-Hallucinator: Mitigating LLM Hallucinations in Code Generation Tasks via Iterative Grounding
von: Eghbali, Aryaz, et al.
Veröffentlicht: (2024)
von: Eghbali, Aryaz, et al.
Veröffentlicht: (2024)
Artisan: Agentic Artifact Evaluation
von: Baek, Doehyun, et al.
Veröffentlicht: (2026)
von: Baek, Doehyun, et al.
Veröffentlicht: (2026)
Defects4C: Benchmarking Large Language Model Repair Capability with C/C++ Bugs
von: Wang, Jian, et al.
Veröffentlicht: (2025)
von: Wang, Jian, et al.
Veröffentlicht: (2025)
GitBug-Java: A Reproducible Benchmark of Recent Java Bugs
von: Silva, André, et al.
Veröffentlicht: (2024)
von: Silva, André, et al.
Veröffentlicht: (2024)
Agentic AI Software Engineers: Programming with Trust
von: Roychoudhury, Abhik, et al.
Veröffentlicht: (2025)
von: Roychoudhury, Abhik, et al.
Veröffentlicht: (2025)
NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition
von: Deng, Le, et al.
Veröffentlicht: (2025)
von: Deng, Le, et al.
Veröffentlicht: (2025)
LLM4FP: LLM-Based Program Generation for Triggering Floating-Point Inconsistencies Across Compilers
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
HotBugs.jar: A Benchmark of Hot Fixes for Time-Critical Bugs
von: Hanna, Carol, et al.
Veröffentlicht: (2025)
von: Hanna, Carol, et al.
Veröffentlicht: (2025)
HaPy-Bug -- Human Annotated Python Bug Resolution Dataset
von: Przymus, Piotr, et al.
Veröffentlicht: (2025)
von: Przymus, Piotr, et al.
Veröffentlicht: (2025)
Black-Box Bug-Amplification for Multithreaded Software
von: Weiss, Yeshayahu, et al.
Veröffentlicht: (2025)
von: Weiss, Yeshayahu, et al.
Veröffentlicht: (2025)
An Empirical Study of Interaction Bugs in ROS-based Software
von: Chen, Zhixiang, et al.
Veröffentlicht: (2025)
von: Chen, Zhixiang, et al.
Veröffentlicht: (2025)
GitBug-Actions: Building Reproducible Bug-Fix Benchmarks with GitHub Actions
von: Saavedra, Nuno, et al.
Veröffentlicht: (2023)
von: Saavedra, Nuno, et al.
Veröffentlicht: (2023)
A Comprehensive Survey on Benchmarks and Solutions in Software Engineering of LLM-Empowered Agentic System
von: Guo, Jiale, et al.
Veröffentlicht: (2025)
von: Guo, Jiale, et al.
Veröffentlicht: (2025)
Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study
von: Wang, You, et al.
Veröffentlicht: (2025)
von: Wang, You, et al.
Veröffentlicht: (2025)
RippleGUItester: Change-Aware Exploratory Testing
von: Su, Yanqi, et al.
Veröffentlicht: (2026)
von: Su, Yanqi, et al.
Veröffentlicht: (2026)
ChangeGuard: Validating Code Changes via Pairwise Learning-Guided Execution
von: Gröninger, Lars, et al.
Veröffentlicht: (2024)
von: Gröninger, Lars, et al.
Veröffentlicht: (2024)
Names Are All You Need: Effective and Safe Regression Test Selection for Python
von: Wang, You, et al.
Veröffentlicht: (2026)
von: Wang, You, et al.
Veröffentlicht: (2026)
Bug Priority Change Prediction: An Exploratory Study on Apache Software
von: Cai, Guangzong, et al.
Veröffentlicht: (2025)
von: Cai, Guangzong, et al.
Veröffentlicht: (2025)
Isolating Compiler Bugs through Compilation Steps Analysis
von: Liu, Yujie, et al.
Veröffentlicht: (2025)
von: Liu, Yujie, et al.
Veröffentlicht: (2025)
Treefix: Enabling Execution with a Tree of Prefixes
von: Souza, Beatriz, et al.
Veröffentlicht: (2025)
von: Souza, Beatriz, et al.
Veröffentlicht: (2025)
You Name It, I Run It: An LLM Agent to Execute Tests of Arbitrary Projects
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
A Comprehensive Study of Bugs in Modern Distributed Deep Learning Systems
von: Ma, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Ma, Xiaoxue, et al.
Veröffentlicht: (2025)
Requirements-Based Test Generation: A Comprehensive Survey
von: Yang, Zhenzhen, et al.
Veröffentlicht: (2025)
von: Yang, Zhenzhen, et al.
Veröffentlicht: (2025)
Automatically Detecting Heterogeneous Bugs in High-Performance Computing Scientific Software
von: Davis, Matthew, et al.
Veröffentlicht: (2025)
von: Davis, Matthew, et al.
Veröffentlicht: (2025)
Explaining Software Bugs Leveraging Code Structures in Neural Machine Translation
von: Mahbub, Parvez, et al.
Veröffentlicht: (2022)
von: Mahbub, Parvez, et al.
Veröffentlicht: (2022)
Analyzing Quantum Programs with LintQ: A Static Analysis Framework for Qiskit
von: Paltenghi, Matteo, et al.
Veröffentlicht: (2023)
von: Paltenghi, Matteo, et al.
Veröffentlicht: (2023)
Characterizing Bugs and Quality Attributes in Quantum Software: A Large-Scale Empirical Study
von: Yousuf, Mir Mohammad, et al.
Veröffentlicht: (2025)
von: Yousuf, Mir Mohammad, et al.
Veröffentlicht: (2025)
A Deep Dive into Large Language Models for Automated Bug Localization and Repair
von: Hossain, Soneya Binta, et al.
Veröffentlicht: (2024)
von: Hossain, Soneya Binta, et al.
Veröffentlicht: (2024)
Refactoring $\neq$ Bug-Inducing: Improving Defect Prediction with Code Change Tactics Analysis
von: Niu, Feifei, et al.
Veröffentlicht: (2025)
von: Niu, Feifei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DyPyBench: A Benchmark of Executable Python Software
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024) -
AgentStepper: Interactive Debugging of Software Development Agents
von: Hutter, Robert, et al.
Veröffentlicht: (2026) -
Evaluating LLM Agents on Automated Software Analysis Tasks
von: Bouzenia, Islem, et al.
Veröffentlicht: (2026) -
Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories
von: Bouzenia, Islem, et al.
Veröffentlicht: (2025) -
BioDefect: The First Dataset for Defect Detection in Bioinformatics Software
von: Xu, Tianxiang, et al.
Veröffentlicht: (2026)