Efficient Uncertainty in LLMs through Evidential Knowledge Distillation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nemani, Lakshmana Sri Harsha, Srijith, P. K., Kuśmierczyk, Tomasz
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912499823017984
author Nemani, Lakshmana Sri Harsha
Srijith, P. K.
Kuśmierczyk, Tomasz
author_facet Nemani, Lakshmana Sri Harsha
Srijith, P. K.
Kuśmierczyk, Tomasz
contents Accurate uncertainty quantification remains a key challenge for standard LLMs, prompting the adoption of Bayesian and ensemble-based methods. However, such methods typically necessitate computationally expensive sampling, involving multiple forward passes to effectively estimate predictive uncertainty. In this paper, we introduce a novel approach enabling efficient and effective uncertainty estimation in LLMs without sacrificing performance. Specifically, we distill uncertainty-aware teacher models - originally requiring multiple forward passes - into compact student models sharing the same architecture but fine-tuned using Low-Rank Adaptation (LoRA). We compare two distinct distillation strategies: one in which the student employs traditional softmax-based outputs, and another in which the student leverages Dirichlet-distributed outputs to explicitly model epistemic uncertainty via evidential learning. Empirical evaluations on classification datasets demonstrate that such students can achieve comparable or superior predictive and uncertainty quantification performance relative to their teacher models, while critically requiring only a single forward pass. To our knowledge, this is the first demonstration that immediate and robust uncertainty quantification can be achieved in LLMs through evidential distillation.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18366
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Uncertainty in LLMs through Evidential Knowledge Distillation
Nemani, Lakshmana Sri Harsha
Srijith, P. K.
Kuśmierczyk, Tomasz
Machine Learning
Accurate uncertainty quantification remains a key challenge for standard LLMs, prompting the adoption of Bayesian and ensemble-based methods. However, such methods typically necessitate computationally expensive sampling, involving multiple forward passes to effectively estimate predictive uncertainty. In this paper, we introduce a novel approach enabling efficient and effective uncertainty estimation in LLMs without sacrificing performance. Specifically, we distill uncertainty-aware teacher models - originally requiring multiple forward passes - into compact student models sharing the same architecture but fine-tuned using Low-Rank Adaptation (LoRA). We compare two distinct distillation strategies: one in which the student employs traditional softmax-based outputs, and another in which the student leverages Dirichlet-distributed outputs to explicitly model epistemic uncertainty via evidential learning. Empirical evaluations on classification datasets demonstrate that such students can achieve comparable or superior predictive and uncertainty quantification performance relative to their teacher models, while critically requiring only a single forward pass. To our knowledge, this is the first demonstration that immediate and robust uncertainty quantification can be achieved in LLMs through evidential distillation.
title Efficient Uncertainty in LLMs through Evidential Knowledge Distillation
topic Machine Learning
url https://arxiv.org/abs/2507.18366