An Empirical Evaluation of LLM-Generated Code Security Across Prompting Methods

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kharma, Mohammed, Sabbah, Ahmed, Alkhanafseh, Mohammad, Hammoudeh, Mohammad, Mohaisen, David
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918519441981440
author Kharma, Mohammed
Sabbah, Ahmed
Alkhanafseh, Mohammad
Hammoudeh, Mohammad
Mohaisen, David
author_facet Kharma, Mohammed
Sabbah, Ahmed
Alkhanafseh, Mohammad
Hammoudeh, Mohammad
Mohaisen, David
contents The growing use of Large Language Models (LLMs) for automated code generation has enhanced software development efficiency, but often at the cost of security. Generated code frequently overlooks critical concerns, leaving it vulnerable to issues such as weak encryption and improper input validation. To investigate this problem, we present a comprehensive empirical evaluation of the security quality of LLM-generated code across five LLMs and four programming languages (Java, C++, C, and Python), examining the impact of multiple prompt engineering methods. We introduce a weaknesses-aware zero-shot chain-of-thought (WA-0CoT) prompting strategy that enriches prompts with security context using CWE mappings to guide model reasoning. Our empirical analysis, supported by chi-square tests, finds no statistically significant reductions in vulnerability frequency or density across prompt methods. However, prompting strategies, including WA-0CoT, systematically influence the compositional distribution of CWE categories, with effects varying by programming language. These findings suggest that while security-aware prompting alters the structure of generated weaknesses, prompt engineering alone is insufficient to reliably reduce overall vulnerability levels. The results highlight the importance of language-aware and model-aware prompt design when evaluating the security properties of LLM-generated code.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24298
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle An Empirical Evaluation of LLM-Generated Code Security Across Prompting Methods
Kharma, Mohammed
Sabbah, Ahmed
Alkhanafseh, Mohammad
Hammoudeh, Mohammad
Mohaisen, David
Cryptography and Security
Artificial Intelligence
Machine Learning
The growing use of Large Language Models (LLMs) for automated code generation has enhanced software development efficiency, but often at the cost of security. Generated code frequently overlooks critical concerns, leaving it vulnerable to issues such as weak encryption and improper input validation. To investigate this problem, we present a comprehensive empirical evaluation of the security quality of LLM-generated code across five LLMs and four programming languages (Java, C++, C, and Python), examining the impact of multiple prompt engineering methods. We introduce a weaknesses-aware zero-shot chain-of-thought (WA-0CoT) prompting strategy that enriches prompts with security context using CWE mappings to guide model reasoning. Our empirical analysis, supported by chi-square tests, finds no statistically significant reductions in vulnerability frequency or density across prompt methods. However, prompting strategies, including WA-0CoT, systematically influence the compositional distribution of CWE categories, with effects varying by programming language. These findings suggest that while security-aware prompting alters the structure of generated weaknesses, prompt engineering alone is insufficient to reliably reduce overall vulnerability levels. The results highlight the importance of language-aware and model-aware prompt design when evaluating the security properties of LLM-generated code.
title An Empirical Evaluation of LLM-Generated Code Security Across Prompting Methods
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.24298