Real-Time Detection of Hallucinated Entities in Long-Form Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Obeso, Oscar, Arditi, Andy, Ferrando, Javier, Freeman, Joshua, Holmes, Cameron, Nanda, Neel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912878266679296
author Obeso, Oscar
Arditi, Andy
Ferrando, Javier
Freeman, Joshua
Holmes, Cameron
Nanda, Neel
author_facet Obeso, Oscar
Arditi, Andy
Ferrando, Javier
Freeman, Joshua
Holmes, Cameron
Nanda, Neel
contents Large language models are now routinely used in high-stakes applications where hallucinations can cause serious harm, such as medical consultations or legal advice. Existing hallucination detection methods, however, are impractical for real-world use, as they are either limited to short factual queries or require costly external verification. We present a cheap, scalable method for real-time identification of hallucinated tokens in long-form generations, and scale it effectively to 70B parameter models. Our approach targets entity-level hallucinations-e.g., fabricated names, dates, citations-rather than claim-level, thereby naturally mapping to token-level labels and enabling streaming detection. We develop an annotation methodology that leverages web search to annotate model responses with grounded labels indicating which tokens correspond to fabricated entities. This dataset enables us to train effective hallucination classifiers with simple and efficient methods such as linear probes. Evaluating across four model families, our classifiers consistently outperform baselines on long-form responses, including more expensive methods such as semantic entropy (e.g., AUC 0.90 vs 0.71 for Llama-3.3-70B), and are also an improvement in short-form question-answering settings. Despite being trained only to detect hallucinated entities, our probes effectively detect incorrect answers in mathematical reasoning tasks, indicating generalization beyond entities. While our annotation methodology is expensive, we find that annotated responses from one model can be used to train effective classifiers on other models; accordingly, we publicly release our datasets to facilitate reuse. Overall, our work suggests a promising new approach for scalable, real-world hallucination detection.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03531
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Real-Time Detection of Hallucinated Entities in Long-Form Generation
Obeso, Oscar
Arditi, Andy
Ferrando, Javier
Freeman, Joshua
Holmes, Cameron
Nanda, Neel
Computation and Language
Artificial Intelligence
Machine Learning
Large language models are now routinely used in high-stakes applications where hallucinations can cause serious harm, such as medical consultations or legal advice. Existing hallucination detection methods, however, are impractical for real-world use, as they are either limited to short factual queries or require costly external verification. We present a cheap, scalable method for real-time identification of hallucinated tokens in long-form generations, and scale it effectively to 70B parameter models. Our approach targets entity-level hallucinations-e.g., fabricated names, dates, citations-rather than claim-level, thereby naturally mapping to token-level labels and enabling streaming detection. We develop an annotation methodology that leverages web search to annotate model responses with grounded labels indicating which tokens correspond to fabricated entities. This dataset enables us to train effective hallucination classifiers with simple and efficient methods such as linear probes. Evaluating across four model families, our classifiers consistently outperform baselines on long-form responses, including more expensive methods such as semantic entropy (e.g., AUC 0.90 vs 0.71 for Llama-3.3-70B), and are also an improvement in short-form question-answering settings. Despite being trained only to detect hallucinated entities, our probes effectively detect incorrect answers in mathematical reasoning tasks, indicating generalization beyond entities. While our annotation methodology is expensive, we find that annotated responses from one model can be used to train effective classifiers on other models; accordingly, we publicly release our datasets to facilitate reuse. Overall, our work suggests a promising new approach for scalable, real-world hallucination detection.
title Real-Time Detection of Hallucinated Entities in Long-Form Generation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.03531