Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Xiaoou, Chen, Tiejin, Zhang, Dengjia, Wang, Yaqing, Cheng, Lu, Wei, Hua
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910234919829504
author Liu, Xiaoou
Chen, Tiejin
Zhang, Dengjia
Wang, Yaqing
Cheng, Lu
Wei, Hua
author_facet Liu, Xiaoou
Chen, Tiejin
Zhang, Dengjia
Wang, Yaqing
Cheng, Lu
Wei, Hua
contents Large Language Models have achieved strong performance on reasoning tasks with objective answers by generating step-by-step solutions, but diagnosing where a multi-step reasoning trace might fail remains difficult. Confidence estimation offers a diagnostic signal, yet existing methods are restricted to final answers or require internal model access. In this paper, we introduce Stepwise Confidence Attribution (SCA), a framework for closed-source LLMs that assigns step-level confidence based only on generated reasoning traces. SCA applies the Information Bottleneck principle: steps aligning with consensus structures across correct solutions receive high confidence, while deviations are flagged as potentially erroneous. We propose two complementary methods: (1) NIBS, a non-parametric IB approach measuring consistency without graph structures, and (2) GIBS, a graph-based IB model that learns subgraphs through a differentiable mask to capture logical variability. Extensive experiments on mathematical reasoning and multi-hop question answering show that SCA reliably identifies low-confidence steps strongly correlated with reasoning errors. Moreover, using step-level confidence to guide self-correction improves the correction success rate by up to 13.5\% over answer-level feedback.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19228
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution
Liu, Xiaoou
Chen, Tiejin
Zhang, Dengjia
Wang, Yaqing
Cheng, Lu
Wei, Hua
Computation and Language
Artificial Intelligence
Information Theory
Machine Learning
68T50, 68T37, 68Q32
I.2.7; I.2.6; I.2.4
Large Language Models have achieved strong performance on reasoning tasks with objective answers by generating step-by-step solutions, but diagnosing where a multi-step reasoning trace might fail remains difficult. Confidence estimation offers a diagnostic signal, yet existing methods are restricted to final answers or require internal model access. In this paper, we introduce Stepwise Confidence Attribution (SCA), a framework for closed-source LLMs that assigns step-level confidence based only on generated reasoning traces. SCA applies the Information Bottleneck principle: steps aligning with consensus structures across correct solutions receive high confidence, while deviations are flagged as potentially erroneous. We propose two complementary methods: (1) NIBS, a non-parametric IB approach measuring consistency without graph structures, and (2) GIBS, a graph-based IB model that learns subgraphs through a differentiable mask to capture logical variability. Extensive experiments on mathematical reasoning and multi-hop question answering show that SCA reliably identifies low-confidence steps strongly correlated with reasoning errors. Moreover, using step-level confidence to guide self-correction improves the correction success rate by up to 13.5\% over answer-level feedback.
title Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution
topic Computation and Language
Artificial Intelligence
Information Theory
Machine Learning
68T50, 68T37, 68Q32
I.2.7; I.2.6; I.2.4
url https://arxiv.org/abs/2605.19228