WildCode: An Empirical Analysis of Code Generated by ChatGPT

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Khanmohammadi, Kobra, Roy, Pooria, Khoury, Raphael, Hamou-Lhadj, Abdelwahab, Konan, Wilfried Patrick
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915653183602688
author Khanmohammadi, Kobra
Roy, Pooria
Khoury, Raphael
Hamou-Lhadj, Abdelwahab
Konan, Wilfried Patrick
author_facet Khanmohammadi, Kobra
Roy, Pooria
Khoury, Raphael
Hamou-Lhadj, Abdelwahab
Konan, Wilfried Patrick
contents LLM models are increasingly used to generate code, but the quality and security of this code are often uncertain. Several recent studies have raised alarm bells, indicating that such AI-generated code may be particularly vulnerable to cyberattacks. However, most of these studies rely on code that is generated specifically for the study, which raises questions about the realism of such experiments. In this study, we perform a large-scale empirical analysis of real-life code generated by ChatGPT. We evaluate code generated by ChatGPT both with respect to correctness and security and delve into the intentions of users who request code from the model. Our research confirms previous studies that used synthetic queries and yielded evidence that LLM-generated code is often inadequate with respect to security. We also find that users exhibit little curiosity about the security features of the code they ask LLMs to generate, as evidenced by their lack of queries on this topic.
format Preprint
id arxiv_https___arxiv_org_abs_2512_04259
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WildCode: An Empirical Analysis of Code Generated by ChatGPT
Khanmohammadi, Kobra
Roy, Pooria
Khoury, Raphael
Hamou-Lhadj, Abdelwahab
Konan, Wilfried Patrick
Cryptography and Security
Software Engineering
LLM models are increasingly used to generate code, but the quality and security of this code are often uncertain. Several recent studies have raised alarm bells, indicating that such AI-generated code may be particularly vulnerable to cyberattacks. However, most of these studies rely on code that is generated specifically for the study, which raises questions about the realism of such experiments. In this study, we perform a large-scale empirical analysis of real-life code generated by ChatGPT. We evaluate code generated by ChatGPT both with respect to correctness and security and delve into the intentions of users who request code from the model. Our research confirms previous studies that used synthetic queries and yielded evidence that LLM-generated code is often inadequate with respect to security. We also find that users exhibit little curiosity about the security features of the code they ask LLMs to generate, as evidenced by their lack of queries on this topic.
title WildCode: An Empirical Analysis of Code Generated by ChatGPT
topic Cryptography and Security
Software Engineering
url https://arxiv.org/abs/2512.04259