Training-free Truthfulness Detection via Value Vectors in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Runheng, Huang, Heyan, Xiao, Xingchen, Wu, Zhijing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915506766741504
author Liu, Runheng
Huang, Heyan
Xiao, Xingchen
Wu, Zhijing
author_facet Liu, Runheng
Huang, Heyan
Xiao, Xingchen
Wu, Zhijing
contents Large language models often generate factually incorrect outputs, motivating efforts to detect the truthfulness of their content. Most existing approaches rely on training probes over internal activations, but these methods suffer from scalability and generalization issues. A recent training-free method, NoVo, addresses this challenge by exploiting statistical patterns from the model itself. However, it focuses exclusively on attention mechanisms, potentially overlooking the MLP module-a core component of Transformer models known to support factual recall. In this paper, we show that certain value vectors within MLP modules exhibit truthfulness-related statistical patterns. Building on this insight, we propose TruthV, a simple and interpretable training-free method that detects content truthfulness by leveraging these value vectors. On the NoVo benchmark, TruthV significantly outperforms both NoVo and log-likelihood baselines, demonstrating that MLP modules-despite being neglected in prior training-free efforts-encode rich and useful signals for truthfulness detection. These findings offer new insights into how truthfulness is internally represented in LLMs and motivate further research on scalable and interpretable truthfulness detection.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17932
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Training-free Truthfulness Detection via Value Vectors in LLMs
Liu, Runheng
Huang, Heyan
Xiao, Xingchen
Wu, Zhijing
Computation and Language
Large language models often generate factually incorrect outputs, motivating efforts to detect the truthfulness of their content. Most existing approaches rely on training probes over internal activations, but these methods suffer from scalability and generalization issues. A recent training-free method, NoVo, addresses this challenge by exploiting statistical patterns from the model itself. However, it focuses exclusively on attention mechanisms, potentially overlooking the MLP module-a core component of Transformer models known to support factual recall. In this paper, we show that certain value vectors within MLP modules exhibit truthfulness-related statistical patterns. Building on this insight, we propose TruthV, a simple and interpretable training-free method that detects content truthfulness by leveraging these value vectors. On the NoVo benchmark, TruthV significantly outperforms both NoVo and log-likelihood baselines, demonstrating that MLP modules-despite being neglected in prior training-free efforts-encode rich and useful signals for truthfulness detection. These findings offer new insights into how truthfulness is internally represented in LLMs and motivate further research on scalable and interpretable truthfulness detection.
title Training-free Truthfulness Detection via Value Vectors in LLMs
topic Computation and Language
url https://arxiv.org/abs/2509.17932