Mind the Generation Process: Fine-Grained Confidence Estimation During LLM Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Han, Jinyi, Li, Tingyun, Chen, Shisong, Shi, Jie, Wang, Xinyi, Yue, Guanglei, Liang, Jiaqing, Lin, Xin, Wen, Liqian, Chen, Zulong, Xiao, Yanghua
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916902806224896
author Han, Jinyi
Li, Tingyun
Chen, Shisong
Shi, Jie
Wang, Xinyi
Yue, Guanglei
Liang, Jiaqing
Lin, Xin
Wen, Liqian
Chen, Zulong
Xiao, Yanghua
author_facet Han, Jinyi
Li, Tingyun
Chen, Shisong
Shi, Jie
Wang, Xinyi
Yue, Guanglei
Liang, Jiaqing
Lin, Xin
Wen, Liqian
Chen, Zulong
Xiao, Yanghua
contents While large language models (LLMs) have demonstrated remarkable performance across diverse tasks, they fundamentally lack self-awareness and frequently exhibit overconfidence, assigning high confidence scores to incorrect predictions. Accurate confidence estimation is therefore critical for enhancing the trustworthiness and reliability of LLM-generated outputs. However, existing approaches suffer from coarse-grained scoring mechanisms that fail to provide fine-grained, continuous confidence estimates throughout the generation process. To address these limitations, we introduce FineCE, a novel confidence estimation method that delivers accurate, fine-grained confidence scores during text generation. Specifically, we first develop a comprehensive pipeline for constructing training data that effectively captures the underlying probabilistic distribution of LLM responses, and then train a model to predict confidence scores for arbitrary text sequences in a supervised manner. Furthermore, we propose a Backward Confidence Integration (BCI) strategy that leverages information from the subsequent text to enhance confidence estimation for the current sequence during inference. We also introduce three strategies for identifying optimal positions to perform confidence estimation within the generation process. Extensive experiments on multiple benchmark datasets demonstrate that FineCE consistently outperforms existing classical confidence estimation methods. Our code and all baselines used in the paper are available on GitHub.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12040
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mind the Generation Process: Fine-Grained Confidence Estimation During LLM Generation
Han, Jinyi
Li, Tingyun
Chen, Shisong
Shi, Jie
Wang, Xinyi
Yue, Guanglei
Liang, Jiaqing
Lin, Xin
Wen, Liqian
Chen, Zulong
Xiao, Yanghua
Computation and Language
Artificial Intelligence
While large language models (LLMs) have demonstrated remarkable performance across diverse tasks, they fundamentally lack self-awareness and frequently exhibit overconfidence, assigning high confidence scores to incorrect predictions. Accurate confidence estimation is therefore critical for enhancing the trustworthiness and reliability of LLM-generated outputs. However, existing approaches suffer from coarse-grained scoring mechanisms that fail to provide fine-grained, continuous confidence estimates throughout the generation process. To address these limitations, we introduce FineCE, a novel confidence estimation method that delivers accurate, fine-grained confidence scores during text generation. Specifically, we first develop a comprehensive pipeline for constructing training data that effectively captures the underlying probabilistic distribution of LLM responses, and then train a model to predict confidence scores for arbitrary text sequences in a supervised manner. Furthermore, we propose a Backward Confidence Integration (BCI) strategy that leverages information from the subsequent text to enhance confidence estimation for the current sequence during inference. We also introduce three strategies for identifying optimal positions to perform confidence estimation within the generation process. Extensive experiments on multiple benchmark datasets demonstrate that FineCE consistently outperforms existing classical confidence estimation methods. Our code and all baselines used in the paper are available on GitHub.
title Mind the Generation Process: Fine-Grained Confidence Estimation During LLM Generation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.12040