The Trilemma of Truth in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Savcisens, Germans, Eliassi-Rad, Tina
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911266231025664
author Savcisens, Germans
Eliassi-Rad, Tina
author_facet Savcisens, Germans
Eliassi-Rad, Tina
contents The public often attributes human-like qualities to large language models (LLMs) and assumes they "know" certain things. In reality, LLMs encode information retained during training as internal probabilistic knowledge. This study examines existing methods for probing the veracity of that knowledge and identifies several flawed underlying assumptions. To address these flaws, we introduce sAwMIL (Sparse-Aware Multiple-Instance Learning), a multiclass probing framework that combines multiple-instance learning with conformal prediction. sAwMIL leverages internal activations of LLMs to classify statements as true, false, or neither. We evaluate sAwMIL across 16 open-source LLMs, including default and chat-based variants, on three new curated datasets. Our results show that (1) common probing methods fail to provide a reliable and transferable veracity direction and, in some settings, perform worse than zero-shot prompting; (2) truth and falsehood are not encoded symmetrically; and (3) LLMs encode a third type of signal that is distinct from both true and false.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23921
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Trilemma of Truth in Large Language Models
Savcisens, Germans
Eliassi-Rad, Tina
Computation and Language
Machine Learning
68T50
I.2.6; I.2.7; G.3
The public often attributes human-like qualities to large language models (LLMs) and assumes they "know" certain things. In reality, LLMs encode information retained during training as internal probabilistic knowledge. This study examines existing methods for probing the veracity of that knowledge and identifies several flawed underlying assumptions. To address these flaws, we introduce sAwMIL (Sparse-Aware Multiple-Instance Learning), a multiclass probing framework that combines multiple-instance learning with conformal prediction. sAwMIL leverages internal activations of LLMs to classify statements as true, false, or neither. We evaluate sAwMIL across 16 open-source LLMs, including default and chat-based variants, on three new curated datasets. Our results show that (1) common probing methods fail to provide a reliable and transferable veracity direction and, in some settings, perform worse than zero-shot prompting; (2) truth and falsehood are not encoded symmetrically; and (3) LLMs encode a third type of signal that is distinct from both true and false.
title The Trilemma of Truth in Large Language Models
topic Computation and Language
Machine Learning
68T50
I.2.6; I.2.7; G.3
url https://arxiv.org/abs/2506.23921