From Entropy to Calibrated Uncertainty: Training Language Models to Reason About Uncertainty

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jenane, Azza, Walha, Nassim, Kuhn, Lukas, Buettner, Florian
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915839576375296
author Jenane, Azza
Walha, Nassim
Kuhn, Lukas
Buettner, Florian
author_facet Jenane, Azza
Walha, Nassim
Kuhn, Lukas
Buettner, Florian
contents Large Language Models (LLMs) that can express interpretable and calibrated uncertainty are crucial in high-stakes domains. While methods to compute uncertainty post-hoc exist, they are often sampling-based and therefore computationally expensive or lack calibration. We propose a three-stage pipeline to post-train LLMs to efficiently infer calibrated uncertainty estimates for their responses. First, we compute fine-grained entropy-based uncertainty scores on the training data, capturing the distributional variability of model outputs in embedding space. Second, these scores are calibrated via Platt scaling, producing reliable and human-interpretable uncertainty signals. Finally, the target LLM is post-trained via reinforcement learning to align its policy with these calibrated signals through a verifiable reward function. Unlike post-hoc uncertainty estimation methods, our approach provides interpretable and computationally efficient uncertainty estimates at test time. Experiments show that models trained with our pipeline achieve better calibration than baselines and generalize to unseen tasks without further processing, suggesting that they learn a robust uncertainty reasoning behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2603_06317
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Entropy to Calibrated Uncertainty: Training Language Models to Reason About Uncertainty
Jenane, Azza
Walha, Nassim
Kuhn, Lukas
Buettner, Florian
Machine Learning
Artificial Intelligence
Large Language Models (LLMs) that can express interpretable and calibrated uncertainty are crucial in high-stakes domains. While methods to compute uncertainty post-hoc exist, they are often sampling-based and therefore computationally expensive or lack calibration. We propose a three-stage pipeline to post-train LLMs to efficiently infer calibrated uncertainty estimates for their responses. First, we compute fine-grained entropy-based uncertainty scores on the training data, capturing the distributional variability of model outputs in embedding space. Second, these scores are calibrated via Platt scaling, producing reliable and human-interpretable uncertainty signals. Finally, the target LLM is post-trained via reinforcement learning to align its policy with these calibrated signals through a verifiable reward function. Unlike post-hoc uncertainty estimation methods, our approach provides interpretable and computationally efficient uncertainty estimates at test time. Experiments show that models trained with our pipeline achieve better calibration than baselines and generalize to unseen tasks without further processing, suggesting that they learn a robust uncertainty reasoning behavior.
title From Entropy to Calibrated Uncertainty: Training Language Models to Reason About Uncertainty
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.06317