LLMs are Bug Replicators: An Empirical Study on LLMs' Capability in Completing Bug-prone Code
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Liwei, Ye, Sixiang, Sun, Zeyu, Chen, Xiang, Zhang, Yuxia, Wang, Bo, Zhang, Jie M., Li, Zheng, Liu, Yong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Empirical Study on the Capability of LLMs in Decomposing Bug Reports
by: Chen, Zhiyuan, et al.
Published: (2025)
by: Chen, Zhiyuan, et al.
Published: (2025)
Reducing False Positives in Static Bug Detection with LLMs: An Empirical Study in Industry
by: Du, Xueying, et al.
Published: (2026)
by: Du, Xueying, et al.
Published: (2026)
Prompt Alchemy: Automatic Prompt Refinement for Enhancing Code Generation
by: Ye, Sixiang, et al.
Published: (2025)
by: Ye, Sixiang, et al.
Published: (2025)
An Empirical Study of Refactoring Engine Bugs
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
An Empirical Study of Interaction Bugs in ROS-based Software
by: Chen, Zhixiang, et al.
Published: (2025)
by: Chen, Zhixiang, et al.
Published: (2025)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
by: Vulićević, Jelena Ilić
Published: (2026)
by: Vulićević, Jelena Ilić
Published: (2026)
Can We Enhance Bug Report Quality Using LLMs?: An Empirical Study of LLM-Based Bug Report Generation
by: Acharya, Jagrit, et al.
Published: (2025)
by: Acharya, Jagrit, et al.
Published: (2025)
Can GPT-O1 Kill All Bugs? An Evaluation of GPT-Family LLMs on QuixBugs
by: Hu, Haichuan, et al.
Published: (2024)
by: Hu, Haichuan, et al.
Published: (2024)
An Extensive Replication Study of the ABLoTS Approach for Bug Localization
by: Niu, Feifei, et al.
Published: (2026)
by: Niu, Feifei, et al.
Published: (2026)
BugScope: Learn to Find Bugs Like Human
by: Guo, Jinyao, et al.
Published: (2025)
by: Guo, Jinyao, et al.
Published: (2025)
An Empirical Study of Bugs in Modern LLM Agent Frameworks
by: Zhu, Xinxue, et al.
Published: (2026)
by: Zhu, Xinxue, et al.
Published: (2026)
One Bug, Hundreds Behind: LLMs for Large-Scale Bug Discovery
by: Wu, Qiushi, et al.
Published: (2025)
by: Wu, Qiushi, et al.
Published: (2025)
An Empirical Study on the Potential of LLMs in Automated Software Refactoring
by: Liu, Bo, et al.
Published: (2024)
by: Liu, Bo, et al.
Published: (2024)
Engineering Pitfalls in AI Coding Tools: An Empirical Study of Bugs in Claude Code, Codex, and Gemini CLI
by: Zhang, Ruixin, et al.
Published: (2026)
by: Zhang, Ruixin, et al.
Published: (2026)
Coffee: Boost Your Code LLMs by Fixing Bugs with Feedback
by: Moon, Seungjun, et al.
Published: (2023)
by: Moon, Seungjun, et al.
Published: (2023)
An Empirical Study on Leveraging Images in Automated Bug Report Reproduction
by: Wang, Dingbang, et al.
Published: (2025)
by: Wang, Dingbang, et al.
Published: (2025)
Bugs in Large Language Models Generated Code: An Empirical Study
by: Tambon, Florian, et al.
Published: (2024)
by: Tambon, Florian, et al.
Published: (2024)
Coding in a Bubble? Evaluating LLMs in Resolving Context Adaptation Bugs During Code Adaptation
by: Zhang, Tanghaoran, et al.
Published: (2026)
by: Zhang, Tanghaoran, et al.
Published: (2026)
Characterizing Bugs in Login Processes of Android Applications: An Empirical Study
by: Zhou, Zixu, et al.
Published: (2025)
by: Zhou, Zixu, et al.
Published: (2025)
An Empirical Study on Embodied Artificial Intelligence Robot (EAIR) Software Bugs
by: Liao, Zeqin, et al.
Published: (2025)
by: Liao, Zeqin, et al.
Published: (2025)
Dissecting Bug Triggers and Failure Modes in Modern Agentic Frameworks: An Empirical Study
by: Zhang, Xiaowen, et al.
Published: (2026)
by: Zhang, Xiaowen, et al.
Published: (2026)
BugForge: Constructing and Utilizing DBMS Bug Repository to Enhance DBMS Testing
by: Li, Dawei, et al.
Published: (2026)
by: Li, Dawei, et al.
Published: (2026)
Exploring the Jupyter Ecosystem: An Empirical Study of Bugs and Vulnerabilities
by: Jiang, Wenyuan, et al.
Published: (2025)
by: Jiang, Wenyuan, et al.
Published: (2025)
Bug Priority Change: An Empirical Study on Apache Projects
by: Li, Zengyang, et al.
Published: (2024)
by: Li, Zengyang, et al.
Published: (2024)
An Empirical Study on the Classification of Bug Reports with Machine Learning
by: Andrade, Renato, et al.
Published: (2025)
by: Andrade, Renato, et al.
Published: (2025)
An Empirical Study on Bug Severity Estimation using Source Code Metrics and Static Analysis
by: Mashhadi, Ehsan, et al.
Published: (2022)
by: Mashhadi, Ehsan, et al.
Published: (2022)
An Empirical Study on the Characteristics of Database Access Bugs in Java Applications
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
Understanding Bug-Reproducing Tests: A First Empirical Study
by: Hora, Andre, et al.
Published: (2026)
by: Hora, Andre, et al.
Published: (2026)
Understanding Bugs in Quantum Simulators: An Empirical Study
by: Upadhyay, Krishna, et al.
Published: (2026)
by: Upadhyay, Krishna, et al.
Published: (2026)
On the Impact of Code Comments for Automated Bug-Fixing: An Empirical Study
by: Vitale, Antonio, et al.
Published: (2026)
by: Vitale, Antonio, et al.
Published: (2026)
LLMs as Evaluators: A Novel Approach to Evaluate Bug Report Summarization
by: Kumar, Abhishek, et al.
Published: (2024)
by: Kumar, Abhishek, et al.
Published: (2024)
PreciseBugCollector: Extensible, Executable and Precise Bug-fix Collection
by: Ye, He, et al.
Published: (2023)
by: Ye, He, et al.
Published: (2023)
Bug In the Code Stack: Can LLMs Find Bugs in Large Python Code Stacks
by: Lee, Hokyung, et al.
Published: (2024)
by: Lee, Hokyung, et al.
Published: (2024)
Bug Whispering: Towards Audio Bug Reporting
by: Masserini, Elena, et al.
Published: (2025)
by: Masserini, Elena, et al.
Published: (2025)
ImproBR: Bug Report Improver Using LLMs
by: Akyol, Emre Furkan, et al.
Published: (2026)
by: Akyol, Emre Furkan, et al.
Published: (2026)
From Logic to Toolchains: An Empirical Study of Bugs in the TypeScript Ecosystem
by: Tang, TianYi, et al.
Published: (2026)
by: Tang, TianYi, et al.
Published: (2026)
Does Programming Language Matter? An Empirical Study of Fuzzing Bug Detection
by: Shirai, Tatsuya, et al.
Published: (2026)
by: Shirai, Tatsuya, et al.
Published: (2026)
A Study of Using Multimodal LLMs for Non-Crash Functional Bug Detection in Android Apps
by: Ju, Bangyan, et al.
Published: (2024)
by: Ju, Bangyan, et al.
Published: (2024)
BugRepro: Enhancing Android Bug Reproduction with Domain-Specific Knowledge Integration
by: Yin, Hongrong, et al.
Published: (2025)
by: Yin, Hongrong, et al.
Published: (2025)
From Reviewers' Lens: Understanding Bug Bounty Report Invalid Reasons with LLMs
by: Zheng, Jiangrui, et al.
Published: (2025)
by: Zheng, Jiangrui, et al.
Published: (2025)
Similar Items
-
An Empirical Study on the Capability of LLMs in Decomposing Bug Reports
by: Chen, Zhiyuan, et al.
Published: (2025) -
Reducing False Positives in Static Bug Detection with LLMs: An Empirical Study in Industry
by: Du, Xueying, et al.
Published: (2026) -
Prompt Alchemy: Automatic Prompt Refinement for Enhancing Code Generation
by: Ye, Sixiang, et al.
Published: (2025) -
An Empirical Study of Refactoring Engine Bugs
by: Wang, Haibo, et al.
Published: (2024) -
An Empirical Study of Interaction Bugs in ROS-based Software
by: Chen, Zhixiang, et al.
Published: (2025)