Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
Fuente:
arXiv
Saved in:
| Main Authors: | Matton, Katie, Ness, Robert Osazuwa, Guttag, John, Kıcıman, Emre |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
by: Kıcıman, Emre, et al.
Published: (2023)
by: Kıcıman, Emre, et al.
Published: (2023)
FaithLM: Towards Faithful Explanations for Large Language Models
by: Chuang, Yu-Neng, et al.
Published: (2024)
by: Chuang, Yu-Neng, et al.
Published: (2024)
Improving Domain Generalization in Contrastive Learning using Adaptive Temperature Control
by: Lewis, Robert, et al.
Published: (2026)
by: Lewis, Robert, et al.
Published: (2026)
MedFuzz: Exploring the Robustness of Large Language Models in Medical Question Answering
by: Ness, Robert Osazuwa, et al.
Published: (2024)
by: Ness, Robert Osazuwa, et al.
Published: (2024)
Do Large Language Models Walk Their Talk? Measuring the Gap Between Implicit Associations, Self-Report, and Behavioral Altruism
by: Andric, Sandro
Published: (2025)
by: Andric, Sandro
Published: (2025)
A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia
by: Monea, Giovanni, et al.
Published: (2023)
by: Monea, Giovanni, et al.
Published: (2023)
Still "Talking About Large Language Models": Some Clarifications
by: Shanahan, Murray
Published: (2024)
by: Shanahan, Murray
Published: (2024)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
by: Anand, Nikhil, et al.
Published: (2026)
by: Anand, Nikhil, et al.
Published: (2026)
CELL your Model: Contrastive Explanations for Large Language Models
by: Luss, Ronny, et al.
Published: (2024)
by: Luss, Ronny, et al.
Published: (2024)
Talking Nonsense: Probing Large Language Models' Understanding of Adversarial Gibberish Inputs
by: Cherepanova, Valeriia, et al.
Published: (2024)
by: Cherepanova, Valeriia, et al.
Published: (2024)
Bridging the Bosphorus: Advancing Turkish Large Language Models through Strategies for Low-Resource Language Adaptation and Benchmarking
by: Acikgoz, Emre Can, et al.
Published: (2024)
by: Acikgoz, Emre Can, et al.
Published: (2024)
On Measuring Faithfulness or Self-consistency of Natural Language Explanations
by: Parcalabescu, Letitia, et al.
Published: (2023)
by: Parcalabescu, Letitia, et al.
Published: (2023)
Listen to the Context: Towards Faithful Large Language Models for Retrieval Augmented Generation on Climate Questions
by: Thulke, David, et al.
Published: (2025)
by: Thulke, David, et al.
Published: (2025)
FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
by: Ming, Yifei, et al.
Published: (2024)
by: Ming, Yifei, et al.
Published: (2024)
A Data-Centric Approach To Generate Faithful and High Quality Patient Summaries with Large Language Models
by: Hegselmann, Stefan, et al.
Published: (2024)
by: Hegselmann, Stefan, et al.
Published: (2024)
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks
by: Xu, Yifei, et al.
Published: (2025)
by: Xu, Yifei, et al.
Published: (2025)
Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
by: Meek, Austin, et al.
Published: (2025)
by: Meek, Austin, et al.
Published: (2025)
LLMExplainer: Large Language Model based Bayesian Inference for Graph Explanation Generation
by: Zhang, Jiaxing, et al.
Published: (2024)
by: Zhang, Jiaxing, et al.
Published: (2024)
LLMCheckup: Conversational Examination of Large Language Models via Interpretability Tools and Self-Explanations
by: Wang, Qianli, et al.
Published: (2024)
by: Wang, Qianli, et al.
Published: (2024)
Measuring Social Norms of Large Language Models
by: Yuan, Ye, et al.
Published: (2024)
by: Yuan, Ye, et al.
Published: (2024)
Prompt-Response Semantic Divergence Metrics for Faithfulness Hallucination and Misalignment Detection in Large Language Models
by: Halperin, Igor
Published: (2025)
by: Halperin, Igor
Published: (2025)
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
by: Wang, Qianli, et al.
Published: (2026)
by: Wang, Qianli, et al.
Published: (2026)
Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation
by: Young, Richard J.
Published: (2026)
by: Young, Richard J.
Published: (2026)
Large Language Models Are Better Logical Fallacy Reasoners with Counterargument, Explanation, and Goal-Aware Prompt Formulation
by: Jeong, Jiwon, et al.
Published: (2025)
by: Jeong, Jiwon, et al.
Published: (2025)
Hippocrates: An Open-Source Framework for Advancing Large Language Models in Healthcare
by: Acikgoz, Emre Can, et al.
Published: (2024)
by: Acikgoz, Emre Can, et al.
Published: (2024)
Modeling the Data-Generating Process is Necessary for Out-of-Distribution Generalization
by: Kaur, Jivat Neet, et al.
Published: (2022)
by: Kaur, Jivat Neet, et al.
Published: (2022)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
by: Arcuschin, Iván, et al.
Published: (2025)
by: Arcuschin, Iván, et al.
Published: (2025)
Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning
by: Chen, Yanda, et al.
Published: (2024)
by: Chen, Yanda, et al.
Published: (2024)
Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
by: Yi, Jingwei, et al.
Published: (2023)
by: Yi, Jingwei, et al.
Published: (2023)
Towards Trustable Language Models: Investigating Information Quality of Large Language Models
by: Rejeleene, Rick, et al.
Published: (2024)
by: Rejeleene, Rick, et al.
Published: (2024)
SibylSense: Adaptive Rubric Learning via Memory Tuning and Adversarial Probing
by: Xu, Yifei, et al.
Published: (2026)
by: Xu, Yifei, et al.
Published: (2026)
Having Beer after Prayer? Measuring Cultural Bias in Large Language Models
by: Naous, Tarek, et al.
Published: (2023)
by: Naous, Tarek, et al.
Published: (2023)
Show Me How It's Done: The Role of Explanations in Fine-Tuning Language Models
by: Ballout, Mohamad, et al.
Published: (2024)
by: Ballout, Mohamad, et al.
Published: (2024)
From Faithfulness to Correctness: Generative Reward Models that Think Critically
by: Ma, Qiyao, et al.
Published: (2025)
by: Ma, Qiyao, et al.
Published: (2025)
FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies
by: Cho, Seonglae, et al.
Published: (2025)
by: Cho, Seonglae, et al.
Published: (2025)
Unfamiliar Finetuning Examples Control How Language Models Hallucinate
by: Kang, Katie, et al.
Published: (2024)
by: Kang, Katie, et al.
Published: (2024)
Transformer Circuit Faithfulness Metrics are not Robust
by: Miller, Joseph, et al.
Published: (2024)
by: Miller, Joseph, et al.
Published: (2024)
Measuring the (Un)Faithfulness of Concept-Based Explanations
by: Kumar, Shubham, et al.
Published: (2025)
by: Kumar, Shubham, et al.
Published: (2025)
(Im)possibility of Automated Hallucination Detection in Large Language Models
by: Karbasi, Amin, et al.
Published: (2025)
by: Karbasi, Amin, et al.
Published: (2025)
Group Preference Optimization: Few-Shot Alignment of Large Language Models
by: Zhao, Siyan, et al.
Published: (2023)
by: Zhao, Siyan, et al.
Published: (2023)
Similar Items
-
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
by: Kıcıman, Emre, et al.
Published: (2023) -
FaithLM: Towards Faithful Explanations for Large Language Models
by: Chuang, Yu-Neng, et al.
Published: (2024) -
Improving Domain Generalization in Contrastive Learning using Adaptive Temperature Control
by: Lewis, Robert, et al.
Published: (2026) -
MedFuzz: Exploring the Robustness of Large Language Models in Medical Question Answering
by: Ness, Robert Osazuwa, et al.
Published: (2024) -
Do Large Language Models Walk Their Talk? Measuring the Gap Between Implicit Associations, Self-Report, and Behavioral Altruism
by: Andric, Sandro
Published: (2025)