Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cao, Qi, Gambardella, Andrew, Kojima, Takeshi, Matsuo, Yutaka, Iwasawa, Yusuke
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914411800690688
author Cao, Qi
Gambardella, Andrew
Kojima, Takeshi
Matsuo, Yutaka
Iwasawa, Yusuke
author_facet Cao, Qi
Gambardella, Andrew
Kojima, Takeshi
Matsuo, Yutaka
Iwasawa, Yusuke
contents Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks. However, the truthfulness of their outputs is not guaranteed, and their tendency toward overconfidence further limits reliability. Uncertainty quantification offers a promising way to identify potentially unreliable outputs, but most existing methods rely on repeated sampling or auxiliary models, introducing substantial computational overhead. To address these limitations, we propose Semantic Token Clustering (STC), an efficient uncertainty quantification method that leverages the semantic information inherently encoded in LLMs. Specifically, we group tokens into semantically consistent clusters using embedding clustering and prefix matching, and quantify uncertainty based on the probability mass aggregated over the corresponding semantic cluster. Our approach requires only a single generation and does not depend on auxiliary models. Experimental results show that STC achieves performance comparable to state-of-the-art baselines while substantially reducing computational overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2603_20161
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
Cao, Qi
Gambardella, Andrew
Kojima, Takeshi
Matsuo, Yutaka
Iwasawa, Yusuke
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks. However, the truthfulness of their outputs is not guaranteed, and their tendency toward overconfidence further limits reliability. Uncertainty quantification offers a promising way to identify potentially unreliable outputs, but most existing methods rely on repeated sampling or auxiliary models, introducing substantial computational overhead. To address these limitations, we propose Semantic Token Clustering (STC), an efficient uncertainty quantification method that leverages the semantic information inherently encoded in LLMs. Specifically, we group tokens into semantically consistent clusters using embedding clustering and prefix matching, and quantify uncertainty based on the probability mass aggregated over the corresponding semantic cluster. Our approach requires only a single generation and does not depend on auxiliary models. Experimental results show that STC achieves performance comparable to state-of-the-art baselines while substantially reducing computational overhead.
title Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.20161