Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Raoyuan, Köksal, Abdullatif, Modarressi, Ali, Hedderich, Michael A., Schütze, Hinrich
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918039853727744
author Zhao, Raoyuan
Köksal, Abdullatif
Modarressi, Ali
Hedderich, Michael A.
Schütze, Hinrich
author_facet Zhao, Raoyuan
Köksal, Abdullatif
Modarressi, Ali
Hedderich, Michael A.
Schütze, Hinrich
contents The reliability of large language models (LLMs) is greatly compromised by their tendency to hallucinate, underscoring the need for precise identification of knowledge gaps within LLMs. Various methods for probing such gaps exist, ranging from calibration-based to prompting-based methods. To evaluate these probing methods, in this paper, we propose a new process based on using input variations and quantitative metrics. Through this, we expose two dimensions of inconsistency in knowledge gap probing. (1) Intra-method inconsistency: Minimal non-semantic perturbations in prompts lead to considerable variance in detected knowledge gaps within the same probing method; e.g., the simple variation of shuffling answer options can decrease agreement to around 40%. (2) Cross-method inconsistency: Probing methods contradict each other on whether a model knows the answer. Methods are highly inconsistent -- with decision consistency across methods being as low as 7% -- even though the model, dataset, and prompt are all the same. These findings challenge existing probing methods and highlight the urgent need for perturbation-robust probing frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21701
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing
Zhao, Raoyuan
Köksal, Abdullatif
Modarressi, Ali
Hedderich, Michael A.
Schütze, Hinrich
Computation and Language
The reliability of large language models (LLMs) is greatly compromised by their tendency to hallucinate, underscoring the need for precise identification of knowledge gaps within LLMs. Various methods for probing such gaps exist, ranging from calibration-based to prompting-based methods. To evaluate these probing methods, in this paper, we propose a new process based on using input variations and quantitative metrics. Through this, we expose two dimensions of inconsistency in knowledge gap probing. (1) Intra-method inconsistency: Minimal non-semantic perturbations in prompts lead to considerable variance in detected knowledge gaps within the same probing method; e.g., the simple variation of shuffling answer options can decrease agreement to around 40%. (2) Cross-method inconsistency: Probing methods contradict each other on whether a model knows the answer. Methods are highly inconsistent -- with decision consistency across methods being as low as 7% -- even though the model, dataset, and prompt are all the same. These findings challenge existing probing methods and highlight the urgent need for perturbation-robust probing frameworks.
title Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing
topic Computation and Language
url https://arxiv.org/abs/2505.21701