Calibrating LLM Confidence by Probing Perturbed Representation Stability

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Khanmohammadi, Reza, Miahi, Erfan, Mardikoraem, Mehrsa, Kaur, Simerjot, Brugere, Ivan, Smiley, Charese H., Thind, Kundan, Ghassemi, Mohammad M.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912593841487872
author Khanmohammadi, Reza
Miahi, Erfan
Mardikoraem, Mehrsa
Kaur, Simerjot
Brugere, Ivan
Smiley, Charese H.
Thind, Kundan
Ghassemi, Mohammad M.
author_facet Khanmohammadi, Reza
Miahi, Erfan
Mardikoraem, Mehrsa
Kaur, Simerjot
Brugere, Ivan
Smiley, Charese H.
Thind, Kundan
Ghassemi, Mohammad M.
contents Miscalibration in Large Language Models (LLMs) undermines their reliability, highlighting the need for accurate confidence estimation. We introduce CCPS (Calibrating LLM Confidence by Probing Perturbed Representation Stability), a novel method analyzing internal representational stability in LLMs. CCPS applies targeted adversarial perturbations to final hidden states, extracts features reflecting the model's response to these perturbations, and uses a lightweight classifier to predict answer correctness. CCPS was evaluated on LLMs from 8B to 32B parameters (covering Llama, Qwen, and Mistral architectures) using MMLU and MMLU-Pro benchmarks in both multiple-choice and open-ended formats. Our results show that CCPS significantly outperforms current approaches. Across four LLMs and three MMLU variants, CCPS reduces Expected Calibration Error by approximately 55% and Brier score by 21%, while increasing accuracy by 5 percentage points, Area Under the Precision-Recall Curve by 4 percentage points, and Area Under the Receiver Operating Characteristic Curve by 6 percentage points, all relative to the strongest prior method. CCPS delivers an efficient, broadly applicable, and more accurate solution for estimating LLM confidence, thereby improving their trustworthiness.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21772
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Calibrating LLM Confidence by Probing Perturbed Representation Stability
Khanmohammadi, Reza
Miahi, Erfan
Mardikoraem, Mehrsa
Kaur, Simerjot
Brugere, Ivan
Smiley, Charese H.
Thind, Kundan
Ghassemi, Mohammad M.
Computation and Language
Miscalibration in Large Language Models (LLMs) undermines their reliability, highlighting the need for accurate confidence estimation. We introduce CCPS (Calibrating LLM Confidence by Probing Perturbed Representation Stability), a novel method analyzing internal representational stability in LLMs. CCPS applies targeted adversarial perturbations to final hidden states, extracts features reflecting the model's response to these perturbations, and uses a lightweight classifier to predict answer correctness. CCPS was evaluated on LLMs from 8B to 32B parameters (covering Llama, Qwen, and Mistral architectures) using MMLU and MMLU-Pro benchmarks in both multiple-choice and open-ended formats. Our results show that CCPS significantly outperforms current approaches. Across four LLMs and three MMLU variants, CCPS reduces Expected Calibration Error by approximately 55% and Brier score by 21%, while increasing accuracy by 5 percentage points, Area Under the Precision-Recall Curve by 4 percentage points, and Area Under the Receiver Operating Characteristic Curve by 6 percentage points, all relative to the strongest prior method. CCPS delivers an efficient, broadly applicable, and more accurate solution for estimating LLM confidence, thereby improving their trustworthiness.
title Calibrating LLM Confidence by Probing Perturbed Representation Stability
topic Computation and Language
url https://arxiv.org/abs/2505.21772