Engineering Pitfalls in AI Coding Tools: An Empirical Study of Bugs in Claude Code, Codex, and Gemini CLI
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Ruixin, Dai, Wuyang, Pham, Hung Viet, Uddin, Gias, Yang, Jinqiu, Wang, Song |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ABTest: Behavior-Driven Testing for AI Coding Agents
by: Dai, Wuyang, et al.
Published: (2026)
by: Dai, Wuyang, et al.
Published: (2026)
An Empirical Study on Bug Severity Estimation using Source Code Metrics and Static Analysis
by: Mashhadi, Ehsan, et al.
Published: (2022)
by: Mashhadi, Ehsan, et al.
Published: (2022)
LLM Assisted Coding with Metamorphic Specification Mutation Agent
by: Akhond, Mostafijur Rahman, et al.
Published: (2025)
by: Akhond, Mostafijur Rahman, et al.
Published: (2025)
Evaluating the Environmental Impact of using SLMs and Prompt Engineering for Code Generation
by: Mamun, Md Afif Al, et al.
Published: (2026)
by: Mamun, Md Afif Al, et al.
Published: (2026)
PAGENT: Learning to Patch Software Engineering Agents
by: Xue, Haoran, et al.
Published: (2025)
by: Xue, Haoran, et al.
Published: (2025)
On the Use of Agentic Coding Manifests: An Empirical Study of Claude Code
by: Chatlatanagulchai, Worawalan, et al.
Published: (2025)
by: Chatlatanagulchai, Worawalan, et al.
Published: (2025)
BLAgent: Agentic RAG for File-Level Bug Localization
by: Mamun, Md Afif Al, et al.
Published: (2026)
by: Mamun, Md Afif Al, et al.
Published: (2026)
Checker Bug Detection and Repair in Deep Learning Libraries
by: Harzevili, Nima Shiri, et al.
Published: (2024)
by: Harzevili, Nima Shiri, et al.
Published: (2024)
Selection of Prompt Engineering Techniques for Code Generation through Predicting Code Complexity
by: Wang, Chung-Yu, et al.
Published: (2024)
by: Wang, Chung-Yu, et al.
Published: (2024)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
by: Aleithan, Reem, et al.
Published: (2024)
by: Aleithan, Reem, et al.
Published: (2024)
Codexity: Secure AI-assisted Code Generation
by: Kim, Sung Yong, et al.
Published: (2024)
by: Kim, Sung Yong, et al.
Published: (2024)
An Empirical Study of Static Analysis Tools for Secure Code Review
by: Charoenwet, Wachiraphan, et al.
Published: (2024)
by: Charoenwet, Wachiraphan, et al.
Published: (2024)
Demystifying Errors in LLM Reasoning Traces: An Empirical Study of Code Execution Simulation
by: Abdollahi, Mohammad, et al.
Published: (2025)
by: Abdollahi, Mohammad, et al.
Published: (2025)
LLMs are Bug Replicators: An Empirical Study on LLMs' Capability in Completing Bug-prone Code
by: Guo, Liwei, et al.
Published: (2025)
by: Guo, Liwei, et al.
Published: (2025)
Optimized Log Parsing with Syntactic Modifications
by: Enan, Nafid, et al.
Published: (2025)
by: Enan, Nafid, et al.
Published: (2025)
A Large-Scale Empirical Study of COVID-19 Contact Tracing Mobile App Reviews
by: Parisa, Sifat Ishmam, et al.
Published: (2024)
by: Parisa, Sifat Ishmam, et al.
Published: (2024)
Semantic Reverse Engineering Legacy Software Applications with ChatGPT, Gemini AI, and Claude AI
by: Mancas, Christian, et al.
Published: (2026)
by: Mancas, Christian, et al.
Published: (2026)
Bugs in Large Language Models Generated Code: An Empirical Study
by: Tambon, Florian, et al.
Published: (2024)
by: Tambon, Florian, et al.
Published: (2024)
A Systematic Mapping Study of Crowd Knowledge Enhanced Software Engineering Research Using Stack Overflow
by: Tanzil, Minaoar, et al.
Published: (2024)
by: Tanzil, Minaoar, et al.
Published: (2024)
CFCEval: Evaluating Security Aspects in Code Generated by Large Language Models
by: Cheng, Cheng, et al.
Published: (2025)
by: Cheng, Cheng, et al.
Published: (2025)
Bias Unveiled: Investigating Social Bias in LLM-Generated Code
by: Ling, Lin, et al.
Published: (2024)
by: Ling, Lin, et al.
Published: (2024)
Tracking the Evolution of Static Code Warnings: the State-of-the-Art and a Better Approach
by: Li, Junjie, et al.
Published: (2022)
by: Li, Junjie, et al.
Published: (2022)
Automated Prompt Engineering for Cost-Effective Code Generation Using Evolutionary Algorithm
by: Taherkhani, Hamed, et al.
Published: (2024)
by: Taherkhani, Hamed, et al.
Published: (2024)
Decoding the Configuration of AI Coding Agents: Insights from Claude Code Projects
by: Santos, Helio Victor F., et al.
Published: (2025)
by: Santos, Helio Victor F., et al.
Published: (2025)
Toward Effective Secure Code Reviews: An Empirical Study of Security-Related Coding Weaknesses
by: Charoenwet, Wachiraphan, et al.
Published: (2023)
by: Charoenwet, Wachiraphan, et al.
Published: (2023)
Assessing the Influence of Toxic and Gender Discriminatory Communication on Perceptible Diversity in OSS Projects
by: Sultana, Sayma, et al.
Published: (2024)
by: Sultana, Sayma, et al.
Published: (2024)
BabelCoder: Agentic Code Translation with Specification Alignment
by: Rabbi, Fazle, et al.
Published: (2025)
by: Rabbi, Fazle, et al.
Published: (2025)
A Multi-Language Perspective on the Robustness of LLM Code Generation
by: Rabbi, Fazle, et al.
Published: (2025)
by: Rabbi, Fazle, et al.
Published: (2025)
On the Impact of Code Comments for Automated Bug-Fixing: An Empirical Study
by: Vitale, Antonio, et al.
Published: (2026)
by: Vitale, Antonio, et al.
Published: (2026)
Specification-Driven Code Translation Powered by Large Language Models: How Far Are We?
by: Saha, Soumit Kanti, et al.
Published: (2024)
by: Saha, Soumit Kanti, et al.
Published: (2024)
TriagerX: Dual Transformers for Bug Triaging Tasks with Content and Interaction Based Rankings
by: Mamun, Md Afif Al, et al.
Published: (2025)
by: Mamun, Md Afif Al, et al.
Published: (2025)
Secure-Instruct: An Automated Pipeline for Synthesizing Instruction-Tuning Datasets Using LLMs for Secure Code Generation
by: Li, Junjie, et al.
Published: (2025)
by: Li, Junjie, et al.
Published: (2025)
Prompt Engineering or Fine-Tuning: An Empirical Assessment of LLMs for Code
by: Shin, Jiho, et al.
Published: (2023)
by: Shin, Jiho, et al.
Published: (2023)
StaAgent: An Agentic Framework for Testing Static Analyzers
by: Nnorom, Elijah, et al.
Published: (2025)
by: Nnorom, Elijah, et al.
Published: (2025)
A Survey of Bugs in AI-Generated Code
by: Gao, Ruofan, et al.
Published: (2025)
by: Gao, Ruofan, et al.
Published: (2025)
Beyond Translation Accuracy: Addressing False Failures in LLM-Based Code Translation
by: Rabbi, Fazle, et al.
Published: (2026)
by: Rabbi, Fazle, et al.
Published: (2026)
Configuring Agentic AI Coding Tools: An Exploratory Study
by: Galster, Matthias, et al.
Published: (2026)
by: Galster, Matthias, et al.
Published: (2026)
Is Vibe Coding the Future? An Empirical Assessment of LLM Generated Codes for Construction Safety
by: Uddin, S M Jamil
Published: (2026)
by: Uddin, S M Jamil
Published: (2026)
Can ChatGPT Support Developers? An Empirical Evaluation of Large Language Models for Code Generation
by: Jin, Kailun, et al.
Published: (2024)
by: Jin, Kailun, et al.
Published: (2024)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
by: Xu, Yisen, et al.
Published: (2026)
by: Xu, Yisen, et al.
Published: (2026)
Similar Items
-
ABTest: Behavior-Driven Testing for AI Coding Agents
by: Dai, Wuyang, et al.
Published: (2026) -
An Empirical Study on Bug Severity Estimation using Source Code Metrics and Static Analysis
by: Mashhadi, Ehsan, et al.
Published: (2022) -
LLM Assisted Coding with Metamorphic Specification Mutation Agent
by: Akhond, Mostafijur Rahman, et al.
Published: (2025) -
Evaluating the Environmental Impact of using SLMs and Prompt Engineering for Code Generation
by: Mamun, Md Afif Al, et al.
Published: (2026) -
PAGENT: Learning to Patch Software Engineering Agents
by: Xue, Haoran, et al.
Published: (2025)