Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bouchard, Dylan, Chauhan, Mohit Singh, Ahmad, Zeya, Ra, Ho-Kyeong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911724839370752
author Bouchard, Dylan
Chauhan, Mohit Singh
Ahmad, Zeya
Ra, Ho-Kyeong
author_facet Bouchard, Dylan
Chauhan, Mohit Singh
Ahmad, Zeya
Ra, Ho-Kyeong
contents Large language models have shown impressive capabilities in code generation, yet they often produce functionally incorrect code. Uncertainty quantification (UQ) methods have emerged as a promising approach for detecting hallucinations in natural language generation, but their effectiveness for code generation tasks remains underexplored. We systematically evaluate how UQ techniques transfer to code generation across three programming languages, five LLMs, and over 1,700 problems. We find that some token-probability-based methods generalize effectively without modification, while sampling-based methods relying on natural language inference (NLI) fail because NLI models cannot distinguish functionally different code, causing most responses to collapse into a single semantic cluster. To address this, we introduce functional equivalence methods, a family of code-specific methods that replace NLI-based semantic equivalence with an LLM-based functional equivalence assessment, including functional entropy, a code-specific analog of semantic entropy. Functional equivalence methods achieve top AUROC in 11 out of 15 model-benchmark combinations and the best calibration across most settings, consistently outperforming both NLI-based counterparts and all other methods evaluated.
format Preprint
id arxiv_https___arxiv_org_abs_2605_28500
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification
Bouchard, Dylan
Chauhan, Mohit Singh
Ahmad, Zeya
Ra, Ho-Kyeong
Computation and Language
Artificial Intelligence
Machine Learning
Large language models have shown impressive capabilities in code generation, yet they often produce functionally incorrect code. Uncertainty quantification (UQ) methods have emerged as a promising approach for detecting hallucinations in natural language generation, but their effectiveness for code generation tasks remains underexplored. We systematically evaluate how UQ techniques transfer to code generation across three programming languages, five LLMs, and over 1,700 problems. We find that some token-probability-based methods generalize effectively without modification, while sampling-based methods relying on natural language inference (NLI) fail because NLI models cannot distinguish functionally different code, causing most responses to collapse into a single semantic cluster. To address this, we introduce functional equivalence methods, a family of code-specific methods that replace NLI-based semantic equivalence with an LLM-based functional equivalence assessment, including functional entropy, a code-specific analog of semantic entropy. Functional equivalence methods achieve top AUROC in 11 out of 15 model-benchmark combinations and the best calibration across most settings, consistently outperforming both NLI-based counterparts and all other methods evaluated.
title Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.28500