HalluLens: LLM Hallucination Benchmark

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Bang, Yejin, Ji, Ziwei, Schelten, Alan, Hartshorn, Anthony, Fowler, Tara, Zhang, Cheng, Cancedda, Nicola, Fung, Pascale
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908336520167424
author Bang, Yejin
Ji, Ziwei
Schelten, Alan
Hartshorn, Anthony
Fowler, Tara
Zhang, Cheng
Cancedda, Nicola
Fung, Pascale
author_facet Bang, Yejin
Ji, Ziwei
Schelten, Alan
Hartshorn, Anthony
Fowler, Tara
Zhang, Cheng
Cancedda, Nicola
Fung, Pascale
contents Large language models (LLMs) often generate responses that deviate from user input or training data, a phenomenon known as "hallucination." These hallucinations undermine user trust and hinder the adoption of generative AI systems. Addressing hallucinations is essential for the advancement of LLMs. This paper introduces a comprehensive hallucination benchmark, incorporating both new extrinsic and existing intrinsic evaluation tasks, built upon clear taxonomy of hallucination. A major challenge in benchmarking hallucinations is the lack of a unified framework due to inconsistent definitions and categorizations. We disentangle LLM hallucination from "factuality," proposing a clear taxonomy that distinguishes between extrinsic and intrinsic hallucinations, to promote consistency and facilitate research. Extrinsic hallucinations, where the generated content is not consistent with the training data, are increasingly important as LLMs evolve. Our benchmark includes dynamic test set generation to mitigate data leakage and ensure robustness against such leakage. We also analyze existing benchmarks, highlighting their limitations and saturation. The work aims to: (1) establish a clear taxonomy of hallucinations, (2) introduce new extrinsic hallucination tasks, with data that can be dynamically regenerated to prevent saturation by leakage, (3) provide a comprehensive analysis of existing benchmarks, distinguishing them from factuality evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2504_17550
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HalluLens: LLM Hallucination Benchmark
Bang, Yejin
Ji, Ziwei
Schelten, Alan
Hartshorn, Anthony
Fowler, Tara
Zhang, Cheng
Cancedda, Nicola
Fung, Pascale
Computation and Language
Artificial Intelligence
Large language models (LLMs) often generate responses that deviate from user input or training data, a phenomenon known as "hallucination." These hallucinations undermine user trust and hinder the adoption of generative AI systems. Addressing hallucinations is essential for the advancement of LLMs. This paper introduces a comprehensive hallucination benchmark, incorporating both new extrinsic and existing intrinsic evaluation tasks, built upon clear taxonomy of hallucination. A major challenge in benchmarking hallucinations is the lack of a unified framework due to inconsistent definitions and categorizations. We disentangle LLM hallucination from "factuality," proposing a clear taxonomy that distinguishes between extrinsic and intrinsic hallucinations, to promote consistency and facilitate research. Extrinsic hallucinations, where the generated content is not consistent with the training data, are increasingly important as LLMs evolve. Our benchmark includes dynamic test set generation to mitigate data leakage and ensure robustness against such leakage. We also analyze existing benchmarks, highlighting their limitations and saturation. The work aims to: (1) establish a clear taxonomy of hallucinations, (2) introduce new extrinsic hallucination tasks, with data that can be dynamically regenerated to prevent saturation by leakage, (3) provide a comprehensive analysis of existing benchmarks, distinguishing them from factuality evaluations.
title HalluLens: LLM Hallucination Benchmark
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.17550