Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917387304960000 |
|---|---|
| author | Xu, Haoming Zhao, Ningyuan Yao, Yunzhi Xu, Weihong Wang, Hongru Deng, Xinle Deng, Shumin Pan, Jeff Z. Chen, Huajun Zhang, Ningyu |
| author_facet | Xu, Haoming Zhao, Ningyuan Yao, Yunzhi Xu, Weihong Wang, Hongru Deng, Xinle Deng, Shumin Pan, Jeff Z. Chen, Huajun Zhang, Ningyu |
| contents | As Large Language Models (LLMs) are increasingly deployed in real-world settings, correctness alone is insufficient. Reliable deployment requires maintaining truthful beliefs under contextual perturbations. Existing evaluations largely rely on point-wise confidence like Self-Consistency, which can mask brittle belief. We show that even facts answered with perfect self-consistency can rapidly collapse under mild contextual interference. To address this gap, we propose Neighbor-Consistency Belief (NCB), a structural measure of belief robustness that evaluates response coherence across a conceptual neighborhood. To validate the efficiency of NCB, we introduce a new cognitive stress-testing protocol that probes outputs stability under contextual interference. Experiments across multiple LLMs show that the performance of high-NCB data is relatively more resistant to interference. Finally, we present Structure-Aware Training (SAT), which optimizes context-invariant belief structure and reduces long-tail knowledge brittleness by approximately 30%. Code is available at https://github.com/zjunlp/belief. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_05905 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency Xu, Haoming Zhao, Ningyuan Yao, Yunzhi Xu, Weihong Wang, Hongru Deng, Xinle Deng, Shumin Pan, Jeff Z. Chen, Huajun Zhang, Ningyu Computation and Language Artificial Intelligence Human-Computer Interaction Machine Learning Multiagent Systems As Large Language Models (LLMs) are increasingly deployed in real-world settings, correctness alone is insufficient. Reliable deployment requires maintaining truthful beliefs under contextual perturbations. Existing evaluations largely rely on point-wise confidence like Self-Consistency, which can mask brittle belief. We show that even facts answered with perfect self-consistency can rapidly collapse under mild contextual interference. To address this gap, we propose Neighbor-Consistency Belief (NCB), a structural measure of belief robustness that evaluates response coherence across a conceptual neighborhood. To validate the efficiency of NCB, we introduce a new cognitive stress-testing protocol that probes outputs stability under contextual interference. Experiments across multiple LLMs show that the performance of high-NCB data is relatively more resistant to interference. Finally, we present Structure-Aware Training (SAT), which optimizes context-invariant belief structure and reduces long-tail knowledge brittleness by approximately 30%. Code is available at https://github.com/zjunlp/belief. |
| title | Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency |
| topic | Computation and Language Artificial Intelligence Human-Computer Interaction Machine Learning Multiagent Systems |
| url | https://arxiv.org/abs/2601.05905 |