Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yongjie, Wang, Yibo, Zhou, Xin, Shen, Zhiqi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916756508901376
author Wang, Yongjie
Wang, Yibo
Zhou, Xin
Shen, Zhiqi
author_facet Wang, Yongjie
Wang, Yibo
Zhou, Xin
Shen, Zhiqi
contents Probing techniques have shown promise in revealing how LLMs encode human-interpretable concepts, particularly when applied to curated datasets. However, the factors governing a dataset's suitability for effective probe training are not well-understood. This study hypothesizes that probe performance on such datasets reflects characteristics of both the LLM's generated responses and its internal feature space. Through quantitative analysis of probe performance and LLM response uncertainty across a series of tasks, we find a strong correlation: improved probe performance consistently corresponds to a reduction in response uncertainty, and vice versa. Subsequently, we delve deeper into this correlation through the lens of feature importance analysis. Our findings indicate that high LLM response variance is associated with a larger set of important features, which poses a greater challenge for probe models and often results in diminished performance. Moreover, leveraging the insights from response uncertainty analysis, we are able to identify concrete examples where LLM representations align with human knowledge across diverse domains, offering additional evidence of interpretable reasoning in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18575
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?
Wang, Yongjie
Wang, Yibo
Zhou, Xin
Shen, Zhiqi
Artificial Intelligence
68T50, 68T35
I.2.0
Probing techniques have shown promise in revealing how LLMs encode human-interpretable concepts, particularly when applied to curated datasets. However, the factors governing a dataset's suitability for effective probe training are not well-understood. This study hypothesizes that probe performance on such datasets reflects characteristics of both the LLM's generated responses and its internal feature space. Through quantitative analysis of probe performance and LLM response uncertainty across a series of tasks, we find a strong correlation: improved probe performance consistently corresponds to a reduction in response uncertainty, and vice versa. Subsequently, we delve deeper into this correlation through the lens of feature importance analysis. Our findings indicate that high LLM response variance is associated with a larger set of important features, which poses a greater challenge for probe models and often results in diminished performance. Moreover, leveraging the insights from response uncertainty analysis, we are able to identify concrete examples where LLM representations align with human knowledge across diverse domains, offering additional evidence of interpretable reasoning in LLMs.
title Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?
topic Artificial Intelligence
68T50, 68T35
I.2.0
url https://arxiv.org/abs/2505.18575