Embedding Perturbation may Better Reflect Intermediate-Step Uncertainty in LLM Reasoning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wen, Qihao, Wang, Jiahao, Nan, Yang, He, Pengfei, Tandon, Ravi, Xu, Han
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917492816871424
author Wen, Qihao
Wang, Jiahao
Nan, Yang
He, Pengfei
Tandon, Ravi
Xu, Han
author_facet Wen, Qihao
Wang, Jiahao
Nan, Yang
He, Pengfei
Tandon, Ravi
Xu, Han
contents Large language Models (LLMs) have achieved significant breakthroughs across diverse domains; however, they can still produce unreliable or misleading outputs. For responsible LLM application, Uncertainty Quantification (UQ) techniques are used to estimate a model's uncertainty about its outputs, indicating the likelihood that those outputs may be problematic. For LLM reasoning tasks, it is essential to estimate the uncertainty not only for the final answer, but also for the intermediate steps of the reasoning, as this can enable more fine-grained and targeted interventions. In this study, we explore what UQ metrics better reflect the LLM's "intermediate uncertainty" during reasoning. Our study reveals that an LLM's incorrect reasoning steps tend to contain tokens which are highly sensitive to the perturbations on the preceding token embeddings, indicating the model's uncertainty among multiple competing continuations. In this way, uncertain (possibly incorrect) intermediate steps can be readily identified using this sensitivity score as guidance in practice. In our experiments, we show such perturbation-based metrics achieve stronger uncertainty quantification performance compared with baselines including probability-based, sampling-based and Bayesian-based methods. Meanwhile, such metrics also enjoy good simplicity and efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02427
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Embedding Perturbation may Better Reflect Intermediate-Step Uncertainty in LLM Reasoning
Wen, Qihao
Wang, Jiahao
Nan, Yang
He, Pengfei
Tandon, Ravi
Xu, Han
Machine Learning
Large language Models (LLMs) have achieved significant breakthroughs across diverse domains; however, they can still produce unreliable or misleading outputs. For responsible LLM application, Uncertainty Quantification (UQ) techniques are used to estimate a model's uncertainty about its outputs, indicating the likelihood that those outputs may be problematic. For LLM reasoning tasks, it is essential to estimate the uncertainty not only for the final answer, but also for the intermediate steps of the reasoning, as this can enable more fine-grained and targeted interventions. In this study, we explore what UQ metrics better reflect the LLM's "intermediate uncertainty" during reasoning. Our study reveals that an LLM's incorrect reasoning steps tend to contain tokens which are highly sensitive to the perturbations on the preceding token embeddings, indicating the model's uncertainty among multiple competing continuations. In this way, uncertain (possibly incorrect) intermediate steps can be readily identified using this sensitivity score as guidance in practice. In our experiments, we show such perturbation-based metrics achieve stronger uncertainty quantification performance compared with baselines including probability-based, sampling-based and Bayesian-based methods. Meanwhile, such metrics also enjoy good simplicity and efficiency.
title Embedding Perturbation may Better Reflect Intermediate-Step Uncertainty in LLM Reasoning
topic Machine Learning
url https://arxiv.org/abs/2602.02427