Can GPT-O1 Kill All Bugs? An Evaluation of GPT-Family LLMs on QuixBugs
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Haichuan, Shang, Ye, Xu, Guolin, He, Congqing, Zhang, Quanjun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
One Bug, Hundreds Behind: LLMs for Large-Scale Bug Discovery
by: Wu, Qiushi, et al.
Published: (2025)
by: Wu, Qiushi, et al.
Published: (2025)
Repair-R1: Better Test Before Repair
by: Hu, Haichuan, et al.
Published: (2025)
by: Hu, Haichuan, et al.
Published: (2025)
Can We Enhance Bug Report Quality Using LLMs?: An Empirical Study of LLM-Based Bug Report Generation
by: Acharya, Jagrit, et al.
Published: (2025)
by: Acharya, Jagrit, et al.
Published: (2025)
Bug Analysis Towards Bug Resolution Time Prediction
by: Ozkan, Hasan Yagiz, et al.
Published: (2024)
by: Ozkan, Hasan Yagiz, et al.
Published: (2024)
ImproBR: Bug Report Improver Using LLMs
by: Akyol, Emre Furkan, et al.
Published: (2026)
by: Akyol, Emre Furkan, et al.
Published: (2026)
Automated Duplicate Bug Report Detection in Large Open Bug Repositories
by: Laney, Clare E., et al.
Published: (2025)
by: Laney, Clare E., et al.
Published: (2025)
Coding in a Bubble? Evaluating LLMs in Resolving Context Adaptation Bugs During Code Adaptation
by: Zhang, Tanghaoran, et al.
Published: (2026)
by: Zhang, Tanghaoran, et al.
Published: (2026)
Bug In the Code Stack: Can LLMs Find Bugs in Large Python Code Stacks
by: Lee, Hokyung, et al.
Published: (2024)
by: Lee, Hokyung, et al.
Published: (2024)
Reducing False Positives in Static Bug Detection with LLMs: An Empirical Study in Industry
by: Du, Xueying, et al.
Published: (2026)
by: Du, Xueying, et al.
Published: (2026)
HLSDebugger: Identification and Correction of Logic Bugs in HLS Code with LLM Solutions
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
BugBlitz-AI: An Intelligent QA Assistant
by: Yao, Yi, et al.
Published: (2024)
by: Yao, Yi, et al.
Published: (2024)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
by: Vulićević, Jelena Ilić
Published: (2026)
by: Vulićević, Jelena Ilić
Published: (2026)
PyResBugs: A Dataset of Residual Python Bugs for Natural Language-Driven Fault Injection
by: Cotroneo, Domenico, et al.
Published: (2025)
by: Cotroneo, Domenico, et al.
Published: (2025)
Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs
by: Al-Kaswan, Ali, et al.
Published: (2026)
by: Al-Kaswan, Ali, et al.
Published: (2026)
Benchmarking Mythos-Linked Bug Rediscovery
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
RLocator: Reinforcement Learning for Bug Localization
by: Chakraborty, Partha, et al.
Published: (2023)
by: Chakraborty, Partha, et al.
Published: (2023)
PerfBench: Can Agents Resolve Real-World Performance Bugs?
by: Garg, Spandan, et al.
Published: (2025)
by: Garg, Spandan, et al.
Published: (2025)
Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair
by: de-Fitero-Dominguez, David, et al.
Published: (2025)
by: de-Fitero-Dominguez, David, et al.
Published: (2025)
BugPilot: Complex Bug Generation for Efficient Learning of SWE Skills
by: Sonwane, Atharv, et al.
Published: (2025)
by: Sonwane, Atharv, et al.
Published: (2025)
SWE-Synth: Synthesizing Verifiable Bug-Fix Data to Enable Large Language Models in Resolving Real-World Bugs
by: Pham, Minh V. T., et al.
Published: (2025)
by: Pham, Minh V. T., et al.
Published: (2025)
A Survey of Bugs in AI-Generated Code
by: Gao, Ruofan, et al.
Published: (2025)
by: Gao, Ruofan, et al.
Published: (2025)
Debug2Fix: Can Interactive Debugging Help Coding Agents Fix More Bugs?
by: Garg, Spandan, et al.
Published: (2026)
by: Garg, Spandan, et al.
Published: (2026)
BugSpotter: Automated Generation of Code Debugging Exercises
by: Pădurean, Victor-Alexandru, et al.
Published: (2024)
by: Pădurean, Victor-Alexandru, et al.
Published: (2024)
BLAgent: Agentic RAG for File-Level Bug Localization
by: Mamun, Md Afif Al, et al.
Published: (2026)
by: Mamun, Md Afif Al, et al.
Published: (2026)
AEGIS: An Agent-based Framework for General Bug Reproduction from Issue Descriptions
by: Wang, Xinchen, et al.
Published: (2024)
by: Wang, Xinchen, et al.
Published: (2024)
Agents in the Sandbox: End-to-End Crash Bug Reproduction for Minecraft
by: Yapağcı, Eray, et al.
Published: (2025)
by: Yapağcı, Eray, et al.
Published: (2025)
An Empirical Study on LLM-based Agents for Automated Bug Fixing
by: Meng, Xiangxin, et al.
Published: (2024)
by: Meng, Xiangxin, et al.
Published: (2024)
RFCAudit: An LLM Agent for Functional Bug Detection in Network Protocols
by: Zheng, Mingwei, et al.
Published: (2025)
by: Zheng, Mingwei, et al.
Published: (2025)
Dynamic Cogeneration of Bug Reproduction Test in Agentic Program Repair
by: Cheng, Runxiang, et al.
Published: (2026)
by: Cheng, Runxiang, et al.
Published: (2026)
Go-Oracle: Automated Test Oracle for Go Concurrency Bugs
by: Tsimpourlas, Foivos, et al.
Published: (2024)
by: Tsimpourlas, Foivos, et al.
Published: (2024)
Past, Present, and Future of Bug Tracking in the Generative AI Era
by: Torun, Utku Boran, et al.
Published: (2025)
by: Torun, Utku Boran, et al.
Published: (2025)
MarsCode Agent: AI-native Automated Bug Fixing
by: Liu, Yizhou, et al.
Published: (2024)
by: Liu, Yizhou, et al.
Published: (2024)
Agentic Bug Reproduction for Effective Automated Program Repair at Google
by: Cheng, Runxiang, et al.
Published: (2025)
by: Cheng, Runxiang, et al.
Published: (2025)
Automated Bug Report Prioritization in Large Open-Source Projects
by: Pierson, Riley, et al.
Published: (2025)
by: Pierson, Riley, et al.
Published: (2025)
Bugs in Large Language Models Generated Code: An Empirical Study
by: Tambon, Florian, et al.
Published: (2024)
by: Tambon, Florian, et al.
Published: (2024)
Agentic Property-Based Testing: Finding Bugs Across the Python Ecosystem
by: Maaz, Muhammad, et al.
Published: (2025)
by: Maaz, Muhammad, et al.
Published: (2025)
Empirical Analysis and Detection of Hallucinations in LLM-Generated Bug Report Summaries
by: Nirujan, Hinduja, et al.
Published: (2026)
by: Nirujan, Hinduja, et al.
Published: (2026)
Faster Configuration Performance Bug Testing with Neural Dual-level Prioritization
by: Ma, Youpeng, et al.
Published: (2025)
by: Ma, Youpeng, et al.
Published: (2025)
Fine-Tuning Code Language Models to Detect Cross-Language Bugs
by: Li, Zengyang, et al.
Published: (2025)
by: Li, Zengyang, et al.
Published: (2025)
Can GPT-4 Replicate Empirical Software Engineering Research?
by: Liang, Jenny T., et al.
Published: (2023)
by: Liang, Jenny T., et al.
Published: (2023)
Similar Items
-
One Bug, Hundreds Behind: LLMs for Large-Scale Bug Discovery
by: Wu, Qiushi, et al.
Published: (2025) -
Repair-R1: Better Test Before Repair
by: Hu, Haichuan, et al.
Published: (2025) -
Can We Enhance Bug Report Quality Using LLMs?: An Empirical Study of LLM-Based Bug Report Generation
by: Acharya, Jagrit, et al.
Published: (2025) -
Bug Analysis Towards Bug Resolution Time Prediction
by: Ozkan, Hasan Yagiz, et al.
Published: (2024) -
ImproBR: Bug Report Improver Using LLMs
by: Akyol, Emre Furkan, et al.
Published: (2026)