LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
Fuente:
arXiv
Saved in:
| Main Authors: | Mayne, Harry, Kearns, Ryan Othniel, Yang, Yushi, Bean, Andrew M., Delaney, Eoin, Russell, Chris, Mahdi, Adam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quantifying construct validity in large language model evaluations
by: Kearns, Ryan Othniel
Published: (2026)
by: Kearns, Ryan Othniel
Published: (2026)
LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation
by: Khouja, Jude, et al.
Published: (2025)
by: Khouja, Jude, et al.
Published: (2025)
Can sparse autoencoders be used to decompose and interpret steering vectors?
by: Mayne, Harry, et al.
Published: (2024)
by: Mayne, Harry, et al.
Published: (2024)
Evaluating Model Explanations without Ground Truth
by: Rawal, Kaivalya, et al.
Published: (2025)
by: Rawal, Kaivalya, et al.
Published: (2025)
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
by: Yang, Yushi, et al.
Published: (2024)
by: Yang, Yushi, et al.
Published: (2024)
Experts Don't Cheat: Learning What You Don't Know By Predicting Pairs
by: Johnson, Daniel D., et al.
Published: (2024)
by: Johnson, Daniel D., et al.
Published: (2024)
Evaluating the Ability of Explanations to Disambiguate Models in a Rashomon Set
by: Rawal, Kaivalya, et al.
Published: (2026)
by: Rawal, Kaivalya, et al.
Published: (2026)
Unsupervised Learning Approaches for Identifying ICU Patient Subgroups: Do Results Generalise?
by: Mayne, Harry, et al.
Published: (2024)
by: Mayne, Harry, et al.
Published: (2024)
A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
by: Mayne, Harry, et al.
Published: (2026)
by: Mayne, Harry, et al.
Published: (2026)
Bayesian Mixture-of-Experts: Towards Making LLMs Know What They Don't Know
by: Li, Albus Yizhuo
Published: (2025)
by: Li, Albus Yizhuo
Published: (2025)
Can AI Assistants Know What They Don't Know?
by: Cheng, Qinyuan, et al.
Published: (2024)
by: Cheng, Qinyuan, et al.
Published: (2024)
Imagining What We Don't Know
by: Samuels, Lisa
Published: (2026)
by: Samuels, Lisa
Published: (2026)
Imagining What We Don't Know
by: Samuels, Lisa
Published: (2026)
by: Samuels, Lisa
Published: (2026)
Large Language Models Must Be Taught to Know What They Don't Know
by: Kapoor, Sanyam, et al.
Published: (2024)
by: Kapoor, Sanyam, et al.
Published: (2024)
Evaluating Fine-Tuning Efficiency of Human-Inspired Learning Strategies in Medical Question Answering
by: Yang, Yushi, et al.
Published: (2024)
by: Yang, Yushi, et al.
Published: (2024)
Fine-Tuned LLMs Know They Don't Know: A Parameter-Efficient Approach to Recovering Honesty
by: Shi, Zeyu, et al.
Published: (2025)
by: Shi, Zeyu, et al.
Published: (2025)
Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
EpiCaR: Knowing What You Don't Know Matters for Better Reasoning in LLMs
by: Yeom, Jewon, et al.
Published: (2026)
by: Yeom, Jewon, et al.
Published: (2026)
Resource-constrained Fairness
by: Goethals, Sofie, et al.
Published: (2024)
by: Goethals, Sofie, et al.
Published: (2024)
Coding Agents Don't Know When to Act
by: Gloaguen, Thibaud, et al.
Published: (2026)
by: Gloaguen, Thibaud, et al.
Published: (2026)
Benchmarking is Broken -- Don't Let AI be its Own Judge
by: Cheng, Zerui, et al.
Published: (2025)
by: Cheng, Zerui, et al.
Published: (2025)
Do Retrieval Augmented Language Models Know When They Don't Know?
by: Zhou, Youchao, et al.
Published: (2025)
by: Zhou, Youchao, et al.
Published: (2025)
Visually Dehallucinative Instruction Generation: Know What You Don't Know
by: Cha, Sungguk, et al.
Published: (2024)
by: Cha, Sungguk, et al.
Published: (2024)
Theory-Grounded Evaluation of Human-Like Fallacy Patterns in LLM Reasoning
by: Richardson, Andrew Keenan, et al.
Published: (2025)
by: Richardson, Andrew Keenan, et al.
Published: (2025)
Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
Why Don't You Know? Evaluating the Impact of Uncertainty Sources on Uncertainty Quantification in LLMs
by: Goloburda, Maiya, et al.
Published: (2026)
by: Goloburda, Maiya, et al.
Published: (2026)
FairImagen: Post-Processing for Bias Mitigation in Text-to-Image Models
by: Fu, Zihao, et al.
Published: (2025)
by: Fu, Zihao, et al.
Published: (2025)
Don't Just Translate, Agitate: Using Large Language Models as Devil's Advocates for AI Explanations
by: Suh, Ashley, et al.
Published: (2025)
by: Suh, Ashley, et al.
Published: (2025)
Don't Explain Noise: Robust Counterfactuals for Randomized Ensembles
by: Forel, Alexandre, et al.
Published: (2022)
by: Forel, Alexandre, et al.
Published: (2022)
When You Don't Know the Answer, Say So
Published: (2024)
Published: (2024)
Knowing You Don't Know: Learning When to Continue Search in Multi-round RAG through Self-Practicing
by: Yang, Diji, et al.
Published: (2025)
by: Yang, Diji, et al.
Published: (2025)
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
by: Mei, Zhiting, et al.
Published: (2025)
by: Mei, Zhiting, et al.
Published: (2025)
Know What You Don't Know: Uncertainty Calibration of Process Reward Models
by: Park, Young-Jin, et al.
Published: (2025)
by: Park, Young-Jin, et al.
Published: (2025)
Know What You Don't Know: Selective Prediction for Early Exit DNNs
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
Large language models can help boost food production, but be mindful of their risks
by: De Clercq, Djavan, et al.
Published: (2024)
by: De Clercq, Djavan, et al.
Published: (2024)
Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
by: Lachenmaier, Clara, et al.
Published: (2025)
by: Lachenmaier, Clara, et al.
Published: (2025)
OxonFair: A Flexible Toolkit for Algorithmic Fairness
by: Delaney, Eoin, et al.
Published: (2024)
by: Delaney, Eoin, et al.
Published: (2024)
Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness
by: Cheang, Chi Seng, et al.
Published: (2025)
by: Cheang, Chi Seng, et al.
Published: (2025)
I Don't Know: Explicit Modeling of Uncertainty with an [IDK] Token
by: Cohen, Roi, et al.
Published: (2024)
by: Cohen, Roi, et al.
Published: (2024)
What Librarians Still Don't Know about Free Software
by: Chudnov, Daniel
Published: (2009)
by: Chudnov, Daniel
Published: (2009)
Similar Items
-
Quantifying construct validity in large language model evaluations
by: Kearns, Ryan Othniel
Published: (2026) -
LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation
by: Khouja, Jude, et al.
Published: (2025) -
Can sparse autoencoders be used to decompose and interpret steering vectors?
by: Mayne, Harry, et al.
Published: (2024) -
Evaluating Model Explanations without Ground Truth
by: Rawal, Kaivalya, et al.
Published: (2025) -
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
by: Yang, Yushi, et al.
Published: (2024)