Hallucination Inspector: A Fact-Checking Judge for API Migration

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Tileria, Marcos, Dash, Santanu Kumar, Pârţachi, Profir-Petru, Barr, Earl T.
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911614063607808
author Tileria, Marcos
Dash, Santanu Kumar
Pârţachi, Profir-Petru
Barr, Earl T.
author_facet Tileria, Marcos
Dash, Santanu Kumar
Pârţachi, Profir-Petru
Barr, Earl T.
contents Large Language Models (LLMs) are increasingly deployed in automated software engineering for tasks such as API migration. While LLMs are able to identify migration patterns, they often make mistakes and fail to produce correct glue code to invoke the new API in place of the old one. We call this issue Scaffolding Hallucination, a failure mode where models generate incorrect calling contexts by inventing Phantom Symbols -- such as imaginary imports, constructors, and constants -- that do not exist in the API specification. In this paper, we show that standard metrics cannot be relied upon to detect these instances of hallucination. We propose Hallucination Inspector, a static analysis tool to detect Scaffolding Hallucination in LLM-generated code. Our approach includes a lightweight evaluation framework that verifies symbols extracted from the abstract syntax tree against a knowledge base derived directly from software documentation for the API. A preliminary evaluation on Android API migrations demonstrates that our approach successfully identifies hallucinations and significantly reduces false positives compared to standard metrics and probabilistic judges
format Preprint
id arxiv_https___arxiv_org_abs_2604_20202
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Hallucination Inspector: A Fact-Checking Judge for API Migration
Tileria, Marcos
Dash, Santanu Kumar
Pârţachi, Profir-Petru
Barr, Earl T.
Software Engineering
Large Language Models (LLMs) are increasingly deployed in automated software engineering for tasks such as API migration. While LLMs are able to identify migration patterns, they often make mistakes and fail to produce correct glue code to invoke the new API in place of the old one. We call this issue Scaffolding Hallucination, a failure mode where models generate incorrect calling contexts by inventing Phantom Symbols -- such as imaginary imports, constructors, and constants -- that do not exist in the API specification. In this paper, we show that standard metrics cannot be relied upon to detect these instances of hallucination. We propose Hallucination Inspector, a static analysis tool to detect Scaffolding Hallucination in LLM-generated code. Our approach includes a lightweight evaluation framework that verifies symbols extracted from the abstract syntax tree against a knowledge base derived directly from software documentation for the API. A preliminary evaluation on Android API migrations demonstrates that our approach successfully identifies hallucinations and significantly reduces false positives compared to standard metrics and probabilistic judges
title Hallucination Inspector: A Fact-Checking Judge for API Migration
topic Software Engineering
url https://arxiv.org/abs/2604.20202