Incoherence as Oracle-less Measure of Error in LLM-Based Code Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Valentin, Thomas, Madadi, Ardi, Sapia, Gaetano, Böhme, Marcel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912760812535808
author Valentin, Thomas
Madadi, Ardi
Sapia, Gaetano
Böhme, Marcel
author_facet Valentin, Thomas
Madadi, Ardi
Sapia, Gaetano
Böhme, Marcel
contents Generating code from a natural language programming task is one of the most successful applications of Large Language Models (LLMs). Yet, the generated program may be buggy. Without an oracle, such as an existing, correct implementation or a formal specification, can we somehow estimate how likely the generated program is correct? In this paper, we propose a measure of incorrectness, called *incoherence*, that can be estimated efficiently in the absence of an oracle and allows us to establish a lower bound on the error, i.e., the probability that the LLM-generated program for that specification is incorrect. In our experiments, our incoherence-based methodology can automatically identify about two-thirds of incorrect programs without reports of false positives for the average task. In fact, *an oracle-based evaluation of LLMs can be reliably replaced by an incoherence-based evaluation*. In particular, we find a very strong agreement between the ranking of LLMs by the number of programs deemed correct via an oracle (pass@1) and the ranking of LLMs by the number of programs deemed correct via incoherence.
format Preprint
id arxiv_https___arxiv_org_abs_2507_00057
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Incoherence as Oracle-less Measure of Error in LLM-Based Code Generation
Valentin, Thomas
Madadi, Ardi
Sapia, Gaetano
Böhme, Marcel
Programming Languages
Artificial Intelligence
Machine Learning
Software Engineering
Generating code from a natural language programming task is one of the most successful applications of Large Language Models (LLMs). Yet, the generated program may be buggy. Without an oracle, such as an existing, correct implementation or a formal specification, can we somehow estimate how likely the generated program is correct? In this paper, we propose a measure of incorrectness, called *incoherence*, that can be estimated efficiently in the absence of an oracle and allows us to establish a lower bound on the error, i.e., the probability that the LLM-generated program for that specification is incorrect. In our experiments, our incoherence-based methodology can automatically identify about two-thirds of incorrect programs without reports of false positives for the average task. In fact, *an oracle-based evaluation of LLMs can be reliably replaced by an incoherence-based evaluation*. In particular, we find a very strong agreement between the ranking of LLMs by the number of programs deemed correct via an oracle (pass@1) and the ranking of LLMs by the number of programs deemed correct via incoherence.
title Incoherence as Oracle-less Measure of Error in LLM-Based Code Generation
topic Programming Languages
Artificial Intelligence
Machine Learning
Software Engineering
url https://arxiv.org/abs/2507.00057