Confidence Estimation for LLMs in Multi-turn Interactions
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911682709684224 |
|---|---|
| author | Zhang, Caiqi Yang, Ruihan Zhu, Xiaochen Li, Chengzu Hu, Tiancheng Dong, Yijiang River Yang, Deqing Collier, Nigel |
| author_facet | Zhang, Caiqi Yang, Ruihan Zhu, Xiaochen Li, Chengzu Hu, Tiancheng Dong, Yijiang River Yang, Deqing Collier, Nigel |
| contents | While confidence estimation is a promising direction for mitigating hallucinations in Large Language Models (LLMs), current research overwhelmingly focuses on single-turn settings. The dynamics of model confidence in multi-turn conversations, where context accumulates and ambiguity is progressively resolved, remain largely unexplored. This work presents the first systematic study of confidence estimation in multi-turn interactions, establishing a formal evaluation framework grounded in two key desiderata: per-turn calibration and monotonicity of confidence as more information becomes available. To facilitate this, we introduce novel metrics, including a length-normalized Expected Calibration Error (InfoECE), and a new "Hinter-Guesser" paradigm for generating controlled evaluation datasets. Our experiments reveal that widely-used confidence techniques struggle with calibration and monotonicity in multi-turn dialogues. In contrast, a novel logit-based probe we introduce, P(Sufficient), proves comparatively more effective, robustly tracking evidence accumulation and distinguishing it from conversational filler. Our work provides a foundational methodology for developing more reliable and trustworthy conversational agents. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_02179 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Confidence Estimation for LLMs in Multi-turn Interactions Zhang, Caiqi Yang, Ruihan Zhu, Xiaochen Li, Chengzu Hu, Tiancheng Dong, Yijiang River Yang, Deqing Collier, Nigel Computation and Language While confidence estimation is a promising direction for mitigating hallucinations in Large Language Models (LLMs), current research overwhelmingly focuses on single-turn settings. The dynamics of model confidence in multi-turn conversations, where context accumulates and ambiguity is progressively resolved, remain largely unexplored. This work presents the first systematic study of confidence estimation in multi-turn interactions, establishing a formal evaluation framework grounded in two key desiderata: per-turn calibration and monotonicity of confidence as more information becomes available. To facilitate this, we introduce novel metrics, including a length-normalized Expected Calibration Error (InfoECE), and a new "Hinter-Guesser" paradigm for generating controlled evaluation datasets. Our experiments reveal that widely-used confidence techniques struggle with calibration and monotonicity in multi-turn dialogues. In contrast, a novel logit-based probe we introduce, P(Sufficient), proves comparatively more effective, robustly tracking evidence accumulation and distinguishing it from conversational filler. Our work provides a foundational methodology for developing more reliable and trustworthy conversational agents. |
| title | Confidence Estimation for LLMs in Multi-turn Interactions |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2601.02179 |