The Case for Repeatable, Open, and Expert-Grounded Hallucination Benchmarks in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Norman, Justin D., Rivera, Michael U., Hughes, D. Alex
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911248849829888
author Norman, Justin D.
Rivera, Michael U.
Hughes, D. Alex
author_facet Norman, Justin D.
Rivera, Michael U.
Hughes, D. Alex
contents Plausible, but inaccurate, tokens in model-generated text are widely believed to be pervasive and problematic for the responsible adoption of language models. Despite this concern, there is little scientific work that attempts to measure the prevalence of language model hallucination in a comprehensive way. In this paper, we argue that language models should be evaluated using repeatable, open, and domain-contextualized hallucination benchmarking. We present a taxonomy of hallucinations alongside a case study that demonstrates that when experts are absent from the early stages of data creation, the resulting hallucination metrics lack validity and practical utility.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17345
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Case for Repeatable, Open, and Expert-Grounded Hallucination Benchmarks in Large Language Models
Norman, Justin D.
Rivera, Michael U.
Hughes, D. Alex
Computation and Language
Plausible, but inaccurate, tokens in model-generated text are widely believed to be pervasive and problematic for the responsible adoption of language models. Despite this concern, there is little scientific work that attempts to measure the prevalence of language model hallucination in a comprehensive way. In this paper, we argue that language models should be evaluated using repeatable, open, and domain-contextualized hallucination benchmarking. We present a taxonomy of hallucinations alongside a case study that demonstrates that when experts are absent from the early stages of data creation, the resulting hallucination metrics lack validity and practical utility.
title The Case for Repeatable, Open, and Expert-Grounded Hallucination Benchmarks in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2505.17345