TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911585463697408 |
|---|---|
| author | Zhang, Tunyu Shi, Haizhou Wang, Yibin Wang, Hengyi He, Xiaoxiao Li, Zhuowei Chen, Haoxian Han, Ligong Xu, Kai Zhang, Huan Metaxas, Dimitris Wang, Hao |
| author_facet | Zhang, Tunyu Shi, Haizhou Wang, Yibin Wang, Hengyi He, Xiaoxiao Li, Zhuowei Chen, Haoxian Han, Ligong Xu, Kai Zhang, Huan Metaxas, Dimitris Wang, Hao |
| contents | While Large Language Models (LLMs) have demonstrated impressive capabilities, their output quality remains inconsistent across various application scenarios, making it difficult to identify trustworthy responses, especially in complex tasks requiring multi-step reasoning. In this paper, we propose a Token-level Uncertainty estimation framework for Reasoning (TokUR) that enables LLMs to self-assess and self-improve their responses in mathematical reasoning. Specifically, we introduce low-rank random weight perturbation during LLM decoding to generate predictive distributions for token-level uncertainty estimation, and we aggregate these uncertainty quantities to capture the semantic uncertainty of generated responses. Experiments on mathematical reasoning datasets of varying difficulty demonstrate that TokUR exhibits a strong correlation with answer correctness and model robustness, and the uncertainty signals produced by TokUR can be leveraged to enhance the model's reasoning performance at test time. These results highlight the effectiveness of TokUR as a principled and scalable approach for improving the reliability and interpretability of LLMs in challenging reasoning tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_11737 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning Zhang, Tunyu Shi, Haizhou Wang, Yibin Wang, Hengyi He, Xiaoxiao Li, Zhuowei Chen, Haoxian Han, Ligong Xu, Kai Zhang, Huan Metaxas, Dimitris Wang, Hao Machine Learning Artificial Intelligence Computation and Language While Large Language Models (LLMs) have demonstrated impressive capabilities, their output quality remains inconsistent across various application scenarios, making it difficult to identify trustworthy responses, especially in complex tasks requiring multi-step reasoning. In this paper, we propose a Token-level Uncertainty estimation framework for Reasoning (TokUR) that enables LLMs to self-assess and self-improve their responses in mathematical reasoning. Specifically, we introduce low-rank random weight perturbation during LLM decoding to generate predictive distributions for token-level uncertainty estimation, and we aggregate these uncertainty quantities to capture the semantic uncertainty of generated responses. Experiments on mathematical reasoning datasets of varying difficulty demonstrate that TokUR exhibits a strong correlation with answer correctness and model robustness, and the uncertainty signals produced by TokUR can be leveraged to enhance the model's reasoning performance at test time. These results highlight the effectiveness of TokUR as a principled and scalable approach for improving the reliability and interpretability of LLMs in challenging reasoning tasks. |
| title | TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2505.11737 |