VeriTrail: Closed-Domain Hallucination Detection with Traceability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Metropolitansky, Dasha, Larson, Jonathan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912933179555840
author Metropolitansky, Dasha
Larson, Jonathan
author_facet Metropolitansky, Dasha
Larson, Jonathan
contents Even when instructed to adhere to source material, language models often generate unsubstantiated content - a phenomenon known as "closed-domain hallucination." This risk is amplified in processes with multiple generative steps (MGS), compared to processes with a single generative step (SGS). However, due to the greater complexity of MGS processes, we argue that detecting hallucinations in their final outputs is necessary but not sufficient: it is equally important to trace where hallucinated content was likely introduced and how faithful content may have been derived from the source material through intermediate outputs. To address this need, we present VeriTrail, the first closed-domain hallucination detection method designed to provide traceability for both MGS and SGS processes. We also introduce the first datasets to include all intermediate outputs as well as human annotations of final outputs' faithfulness for their respective MGS processes. We demonstrate that VeriTrail outperforms baseline methods on both datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21786
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VeriTrail: Closed-Domain Hallucination Detection with Traceability
Metropolitansky, Dasha
Larson, Jonathan
Computation and Language
Artificial Intelligence
Even when instructed to adhere to source material, language models often generate unsubstantiated content - a phenomenon known as "closed-domain hallucination." This risk is amplified in processes with multiple generative steps (MGS), compared to processes with a single generative step (SGS). However, due to the greater complexity of MGS processes, we argue that detecting hallucinations in their final outputs is necessary but not sufficient: it is equally important to trace where hallucinated content was likely introduced and how faithful content may have been derived from the source material through intermediate outputs. To address this need, we present VeriTrail, the first closed-domain hallucination detection method designed to provide traceability for both MGS and SGS processes. We also introduce the first datasets to include all intermediate outputs as well as human annotations of final outputs' faithfulness for their respective MGS processes. We demonstrate that VeriTrail outperforms baseline methods on both datasets.
title VeriTrail: Closed-Domain Hallucination Detection with Traceability
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.21786