Can LLMs Find Bugs in Code? An Evaluation from Beginner Errors to Security Vulnerabilities in Python and C++
Fuente:
arXiv
Saved in:
| Main Authors: | Mhatre, Akshay, Nader, Noujoud, Diehl, Patrick, Gupta, Deepti |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-HPC++: Evaluating LLM-Generated Modern C++ and MPI+OpenMP Codes for Scalable Mandelbrot Set Computation
by: Diehl, Patrick, et al.
Published: (2025)
by: Diehl, Patrick, et al.
Published: (2025)
Evaluating AI-generated code for C++, Fortran, Go, Java, Julia, Matlab, Python, R, and Rust
by: Diehl, Patrick, et al.
Published: (2024)
by: Diehl, Patrick, et al.
Published: (2024)
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
by: Vulićević, Jelena Ilić
Published: (2026)
by: Vulićević, Jelena Ilić
Published: (2026)
LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
by: Diehl, Patrick, et al.
Published: (2025)
by: Diehl, Patrick, et al.
Published: (2025)
scicode-lint: Detecting Methodology Bugs in Scientific Python Code with LLM-Generated Patterns
by: Samsonau, Sergey V.
Published: (2026)
by: Samsonau, Sergey V.
Published: (2026)
Bug In the Code Stack: Can LLMs Find Bugs in Large Python Code Stacks
by: Lee, Hokyung, et al.
Published: (2024)
by: Lee, Hokyung, et al.
Published: (2024)
K-ASTRO: Structure-Aware Adaptation of LLMs for Code Vulnerability Detection
by: Zhang, Yifan, et al.
Published: (2022)
by: Zhang, Yifan, et al.
Published: (2022)
Language Models are Better Bug Detector Through Code-Pair Classification
by: Alrashedy, Kamel, et al.
Published: (2023)
by: Alrashedy, Kamel, et al.
Published: (2023)
Challenging Bug Prediction and Repair Models with Synthetic Bugs
by: Ibrahimzada, Ali Reza, et al.
Published: (2023)
by: Ibrahimzada, Ali Reza, et al.
Published: (2023)
Unsupervised Evaluation of Code LLMs with Round-Trip Correctness
by: Allamanis, Miltiadis, et al.
Published: (2024)
by: Allamanis, Miltiadis, et al.
Published: (2024)
Finding Missed Code Size Optimizations in Compilers using LLMs
by: Italiano, Davide, et al.
Published: (2024)
by: Italiano, Davide, et al.
Published: (2024)
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
by: Thillen, Alex, et al.
Published: (2026)
by: Thillen, Alex, et al.
Published: (2026)
PromSec: Prompt Optimization for Secure Generation of Functional Source Code with Large Language Models (LLMs)
by: Nazzal, Mahmoud, et al.
Published: (2024)
by: Nazzal, Mahmoud, et al.
Published: (2024)
CPP-UT-Bench: Can LLMs Write Complex Unit Tests in C++?
by: Bhargava, Vaishnavi, et al.
Published: (2024)
by: Bhargava, Vaishnavi, et al.
Published: (2024)
Rethinking the Evaluation of Secure Code Generation
by: Dai, Shih-Chieh, et al.
Published: (2025)
by: Dai, Shih-Chieh, et al.
Published: (2025)
SynthFix: Adaptive Neuro-Symbolic Code Vulnerability Repair
by: Zhang, Yifan, et al.
Published: (2026)
by: Zhang, Yifan, et al.
Published: (2026)
BugGen: A Self-Correcting Multi-Agent LLM Pipeline for Realistic RTL Bug Synthesis
by: Jasper, Surya, et al.
Published: (2025)
by: Jasper, Surya, et al.
Published: (2025)
On the Impact of Code Comments for Automated Bug-Fixing: An Empirical Study
by: Vitale, Antonio, et al.
Published: (2026)
by: Vitale, Antonio, et al.
Published: (2026)
Enhancing Large Language Models with Faster Code Preprocessing for Vulnerability Detection
by: Gonçalves, José, et al.
Published: (2025)
by: Gonçalves, José, et al.
Published: (2025)
LLM Critics Help Catch LLM Bugs
by: McAleese, Nat, et al.
Published: (2024)
by: McAleese, Nat, et al.
Published: (2024)
PLMGH: What Matters in PLM-GNN Hybrids for Code Classification and Vulnerability Detection
by: Idrissi, Mohamed Taoufik Kaouthar El, et al.
Published: (2026)
by: Idrissi, Mohamed Taoufik Kaouthar El, et al.
Published: (2026)
Teaching Code Refactoring Using LLMs
by: Khairnar, Anshul, et al.
Published: (2025)
by: Khairnar, Anshul, et al.
Published: (2025)
Towards Verified Code Reasoning by LLMs
by: Sistla, Meghana, et al.
Published: (2025)
by: Sistla, Meghana, et al.
Published: (2025)
Enhancing Software Vulnerability Detection Using Code Property Graphs and Convolutional Neural Networks
by: Saimbhi, Amanpreet Singh
Published: (2025)
by: Saimbhi, Amanpreet Singh
Published: (2025)
Where's the Bug? Attention Probing for Scalable Fault Localization
by: Stein, Adam, et al.
Published: (2025)
by: Stein, Adam, et al.
Published: (2025)
The Limits of Long-Context Reasoning in Automated Bug Fixing
by: Raju, Ravi, et al.
Published: (2026)
by: Raju, Ravi, et al.
Published: (2026)
Evaluating the Use of LLMs for Documentation to Code Traceability
by: Alor, Ebube, et al.
Published: (2025)
by: Alor, Ebube, et al.
Published: (2025)
OSS-Bench: Benchmark Generator for Coding LLMs
by: Jiang, Yuancheng, et al.
Published: (2025)
by: Jiang, Yuancheng, et al.
Published: (2025)
Understanding Robustness of Model Editing in Code LLMs
by: Chhetri, Vinaik, et al.
Published: (2025)
by: Chhetri, Vinaik, et al.
Published: (2025)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
by: Mündler, Niels, et al.
Published: (2024)
by: Mündler, Niels, et al.
Published: (2024)
Can Code Language Models Learn Clarification-Seeking Behaviors?
by: Wu, Jie JW, et al.
Published: (2025)
by: Wu, Jie JW, et al.
Published: (2025)
Towards Enhancing the Reproducibility of Deep Learning Bugs: An Empirical Study
by: Shah, Mehil B., et al.
Published: (2024)
by: Shah, Mehil B., et al.
Published: (2024)
evomap: A Toolbox for Dynamic Mapping in Python
by: Matthe, Maximilian
Published: (2025)
by: Matthe, Maximilian
Published: (2025)
Evaluating and Improving the Robustness of Security Attack Detectors Generated by LLMs
by: Pasini, Samuele, et al.
Published: (2024)
by: Pasini, Samuele, et al.
Published: (2024)
Automated Code-centric Software Vulnerability Assessment: How Far Are We? An Empirical Study in C/C++
by: Nguyen, Anh The, et al.
Published: (2024)
by: Nguyen, Anh The, et al.
Published: (2024)
Assessing the Quality and Security of AI-Generated Code: A Quantitative Analysis
by: Sabra, Abbas, et al.
Published: (2025)
by: Sabra, Abbas, et al.
Published: (2025)
LLM-Powered Test Case Generation for Detecting Bugs in Plausible Programs
by: Liu, Kaibo, et al.
Published: (2024)
by: Liu, Kaibo, et al.
Published: (2024)
From Lab to Reality: A Practical Evaluation of Deep Learning Models and LLMs for Vulnerability Detection
by: Lu, Chaomeng, et al.
Published: (2025)
by: Lu, Chaomeng, et al.
Published: (2025)
You Only Train Once: A Flexible Training Framework for Code Vulnerability Detection Driven by Vul-Vector
by: Tian, Bowen, et al.
Published: (2025)
by: Tian, Bowen, et al.
Published: (2025)
CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation
by: Peng, Jinjun, et al.
Published: (2025)
by: Peng, Jinjun, et al.
Published: (2025)
Similar Items
-
LLM-HPC++: Evaluating LLM-Generated Modern C++ and MPI+OpenMP Codes for Scalable Mandelbrot Set Computation
by: Diehl, Patrick, et al.
Published: (2025) -
Evaluating AI-generated code for C++, Fortran, Go, Java, Julia, Matlab, Python, R, and Rust
by: Diehl, Patrick, et al.
Published: (2024) -
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
by: Vulićević, Jelena Ilić
Published: (2026) -
LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
by: Diehl, Patrick, et al.
Published: (2025) -
scicode-lint: Detecting Methodology Bugs in Scientific Python Code with LLM-Generated Patterns
by: Samsonau, Sergey V.
Published: (2026)