Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Fang, Liu, Yang, Shi, Lin, Yang, Zhen, Zhang, Li, Lian, Xiaoli, Li, Zhongqi, Ma, Yuchi
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912835995435008
author Liu, Fang
Liu, Yang
Shi, Lin
Yang, Zhen
Zhang, Li
Lian, Xiaoli
Li, Zhongqi
Ma, Yuchi
author_facet Liu, Fang
Liu, Yang
Shi, Lin
Yang, Zhen
Zhang, Li
Lian, Xiaoli
Li, Zhongqi
Ma, Yuchi
contents The rise of Large Language Models (LLMs) has significantly advanced various applications on software engineering tasks, particularly in code generation. Despite the promising performance, LLMs are prone to generate hallucinations, which means LLMs might produce outputs that deviate from users' intent, exhibit internal inconsistencies, or misaligned with the real-world knowledge, making the deployment of LLMs potentially risky in a wide range of applications. Existing work mainly focuses on investigating the hallucination in the domain of Natural Language Generation (NLG), leaving a gap in comprehensively understanding the types, causes, and impacts of hallucinations in the context of code generation. To bridge the gap, we conducted a thematic analysis of the LLM-generated code to summarize and categorize the hallucinations, as well as their causes and impacts. Our study established a comprehensive taxonomy of code hallucinations, encompassing 3 primary categories and 12 specific categories. Furthermore, we systematically analyzed the distribution of hallucinations, exploring variations among different LLMs and benchmarks. Moreover, we perform an in-depth analysis on the causes and impacts of various hallucinations, aiming to provide valuable insights into hallucination mitigation. Finally, to enhance the correctness and reliability of LLM-generated code in a lightweight manner, we explore training-free hallucination mitigation approaches by prompt enhancing techniques. We believe our findings will shed light on future research about code hallucination evaluation and mitigation, ultimately paving the way for building more effective and reliable code LLMs in the future. The replication package is available at https://github.com/Lorien1128/code_hallucination
format Preprint
id arxiv_https___arxiv_org_abs_2404_00971
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code
Liu, Fang
Liu, Yang
Shi, Lin
Yang, Zhen
Zhang, Li
Lian, Xiaoli
Li, Zhongqi
Ma, Yuchi
Software Engineering
Artificial Intelligence
The rise of Large Language Models (LLMs) has significantly advanced various applications on software engineering tasks, particularly in code generation. Despite the promising performance, LLMs are prone to generate hallucinations, which means LLMs might produce outputs that deviate from users' intent, exhibit internal inconsistencies, or misaligned with the real-world knowledge, making the deployment of LLMs potentially risky in a wide range of applications. Existing work mainly focuses on investigating the hallucination in the domain of Natural Language Generation (NLG), leaving a gap in comprehensively understanding the types, causes, and impacts of hallucinations in the context of code generation. To bridge the gap, we conducted a thematic analysis of the LLM-generated code to summarize and categorize the hallucinations, as well as their causes and impacts. Our study established a comprehensive taxonomy of code hallucinations, encompassing 3 primary categories and 12 specific categories. Furthermore, we systematically analyzed the distribution of hallucinations, exploring variations among different LLMs and benchmarks. Moreover, we perform an in-depth analysis on the causes and impacts of various hallucinations, aiming to provide valuable insights into hallucination mitigation. Finally, to enhance the correctness and reliability of LLM-generated code in a lightweight manner, we explore training-free hallucination mitigation approaches by prompt enhancing techniques. We believe our findings will shed light on future research about code hallucination evaluation and mitigation, ultimately paving the way for building more effective and reliable code LLMs in the future. The replication package is available at https://github.com/Lorien1128/code_hallucination
title Beyond Functional Correctness: Exploring Hallucinations in LLM-Generated Code
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2404.00971