Hallucinations in Code Change to Natural Language Generation: Prevalence and Evaluation of Detection Metrics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Chunhua, Lin, Hong Yi, Thongtanunam, Patanamon
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912533005205504
author Liu, Chunhua
Lin, Hong Yi
Thongtanunam, Patanamon
author_facet Liu, Chunhua
Lin, Hong Yi
Thongtanunam, Patanamon
contents Language models have shown strong capabilities across a wide range of tasks in software engineering, such as code generation, yet they suffer from hallucinations. While hallucinations have been studied independently in natural language and code generation, their occurrence in tasks involving code changes which have a structurally complex and context-dependent format of code remains largely unexplored. This paper presents the first comprehensive analysis of hallucinations in two critical tasks involving code change to natural language generation: commit message generation and code review comment generation. We quantify the prevalence of hallucinations in recent language models and explore a range of metric-based approaches to automatically detect them. Our findings reveal that approximately 50\% of generated code reviews and 20\% of generated commit messages contain hallucinations. Whilst commonly used metrics are weak detectors on their own, combining multiple metrics substantially improves performance. Notably, model confidence and feature attribution metrics effectively contribute to hallucination detection, showing promise for inference-time detection.\footnote{All code and data will be released upon acceptance.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08661
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hallucinations in Code Change to Natural Language Generation: Prevalence and Evaluation of Detection Metrics
Liu, Chunhua
Lin, Hong Yi
Thongtanunam, Patanamon
Software Engineering
Artificial Intelligence
Language models have shown strong capabilities across a wide range of tasks in software engineering, such as code generation, yet they suffer from hallucinations. While hallucinations have been studied independently in natural language and code generation, their occurrence in tasks involving code changes which have a structurally complex and context-dependent format of code remains largely unexplored. This paper presents the first comprehensive analysis of hallucinations in two critical tasks involving code change to natural language generation: commit message generation and code review comment generation. We quantify the prevalence of hallucinations in recent language models and explore a range of metric-based approaches to automatically detect them. Our findings reveal that approximately 50\% of generated code reviews and 20\% of generated commit messages contain hallucinations. Whilst commonly used metrics are weak detectors on their own, combining multiple metrics substantially improves performance. Notably, model confidence and feature attribution metrics effectively contribute to hallucination detection, showing promise for inference-time detection.\footnote{All code and data will be released upon acceptance.
title Hallucinations in Code Change to Natural Language Generation: Prevalence and Evaluation of Detection Metrics
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2508.08661