Steer LLM Latents for Hallucination Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Seongheon, Du, Xuefeng, Yeh, Min-Hsuan, Wang, Haobo, Li, Yixuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913851770929152
author Park, Seongheon
Du, Xuefeng
Yeh, Min-Hsuan
Wang, Haobo
Li, Yixuan
author_facet Park, Seongheon
Du, Xuefeng
Yeh, Min-Hsuan
Wang, Haobo
Li, Yixuan
contents Hallucinations in LLMs pose a significant concern to their safe deployment in real-world applications. Recent approaches have leveraged the latent space of LLMs for hallucination detection, but their embeddings, optimized for linguistic coherence rather than factual accuracy, often fail to clearly separate truthful and hallucinated content. To this end, we propose the Truthfulness Separator Vector (TSV), a lightweight and flexible steering vector that reshapes the LLM's representation space during inference to enhance the separation between truthful and hallucinated outputs, without altering model parameters. Our two-stage framework first trains TSV on a small set of labeled exemplars to form compact and well-separated clusters. It then augments the exemplar set with unlabeled LLM generations, employing an optimal transport-based algorithm for pseudo-labeling combined with a confidence-based filtering process. Extensive experiments demonstrate that TSV achieves state-of-the-art performance with minimal labeled data, exhibiting strong generalization across datasets and providing a practical solution for real-world LLM applications.
format Preprint
id arxiv_https___arxiv_org_abs_2503_01917
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Steer LLM Latents for Hallucination Detection
Park, Seongheon
Du, Xuefeng
Yeh, Min-Hsuan
Wang, Haobo
Li, Yixuan
Machine Learning
Artificial Intelligence
Computation and Language
Hallucinations in LLMs pose a significant concern to their safe deployment in real-world applications. Recent approaches have leveraged the latent space of LLMs for hallucination detection, but their embeddings, optimized for linguistic coherence rather than factual accuracy, often fail to clearly separate truthful and hallucinated content. To this end, we propose the Truthfulness Separator Vector (TSV), a lightweight and flexible steering vector that reshapes the LLM's representation space during inference to enhance the separation between truthful and hallucinated outputs, without altering model parameters. Our two-stage framework first trains TSV on a small set of labeled exemplars to form compact and well-separated clusters. It then augments the exemplar set with unlabeled LLM generations, employing an optimal transport-based algorithm for pseudo-labeling combined with a confidence-based filtering process. Extensive experiments demonstrate that TSV achieves state-of-the-art performance with minimal labeled data, exhibiting strong generalization across datasets and providing a practical solution for real-world LLM applications.
title Steer LLM Latents for Hallucination Detection
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2503.01917