TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Tunyu, Shi, Haizhou, Wang, Yibin, Wang, Hengyi, He, Xiaoxiao, Li, Zhuowei, Chen, Haoxian, Han, Ligong, Xu, Kai, Zhang, Huan, Metaxas, Dimitris, Wang, Hao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911585463697408
author Zhang, Tunyu
Shi, Haizhou
Wang, Yibin
Wang, Hengyi
He, Xiaoxiao
Li, Zhuowei
Chen, Haoxian
Han, Ligong
Xu, Kai
Zhang, Huan
Metaxas, Dimitris
Wang, Hao
author_facet Zhang, Tunyu
Shi, Haizhou
Wang, Yibin
Wang, Hengyi
He, Xiaoxiao
Li, Zhuowei
Chen, Haoxian
Han, Ligong
Xu, Kai
Zhang, Huan
Metaxas, Dimitris
Wang, Hao
contents While Large Language Models (LLMs) have demonstrated impressive capabilities, their output quality remains inconsistent across various application scenarios, making it difficult to identify trustworthy responses, especially in complex tasks requiring multi-step reasoning. In this paper, we propose a Token-level Uncertainty estimation framework for Reasoning (TokUR) that enables LLMs to self-assess and self-improve their responses in mathematical reasoning. Specifically, we introduce low-rank random weight perturbation during LLM decoding to generate predictive distributions for token-level uncertainty estimation, and we aggregate these uncertainty quantities to capture the semantic uncertainty of generated responses. Experiments on mathematical reasoning datasets of varying difficulty demonstrate that TokUR exhibits a strong correlation with answer correctness and model robustness, and the uncertainty signals produced by TokUR can be leveraged to enhance the model's reasoning performance at test time. These results highlight the effectiveness of TokUR as a principled and scalable approach for improving the reliability and interpretability of LLMs in challenging reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11737
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
Zhang, Tunyu
Shi, Haizhou
Wang, Yibin
Wang, Hengyi
He, Xiaoxiao
Li, Zhuowei
Chen, Haoxian
Han, Ligong
Xu, Kai
Zhang, Huan
Metaxas, Dimitris
Wang, Hao
Machine Learning
Artificial Intelligence
Computation and Language
While Large Language Models (LLMs) have demonstrated impressive capabilities, their output quality remains inconsistent across various application scenarios, making it difficult to identify trustworthy responses, especially in complex tasks requiring multi-step reasoning. In this paper, we propose a Token-level Uncertainty estimation framework for Reasoning (TokUR) that enables LLMs to self-assess and self-improve their responses in mathematical reasoning. Specifically, we introduce low-rank random weight perturbation during LLM decoding to generate predictive distributions for token-level uncertainty estimation, and we aggregate these uncertainty quantities to capture the semantic uncertainty of generated responses. Experiments on mathematical reasoning datasets of varying difficulty demonstrate that TokUR exhibits a strong correlation with answer correctness and model robustness, and the uncertainty signals produced by TokUR can be leveraged to enhance the model's reasoning performance at test time. These results highlight the effectiveness of TokUR as a principled and scalable approach for improving the reliability and interpretability of LLMs in challenging reasoning tasks.
title TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.11737