Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Agarwal, Chirag, Tanneru, Sree Harsha, Lakkaraju, Himabindu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models
by: Tanneru, Sree Harsha, et al.
Published: (2024)
by: Tanneru, Sree Harsha, et al.
Published: (2024)
Understanding the Effects of Iterative Prompting on Truthfulness
by: Krishna, Satyapriya, et al.
Published: (2024)
by: Krishna, Satyapriya, et al.
Published: (2024)
On the Impact of Fine-Tuning on Chain-of-Thought Reasoning
by: Lobo, Elita, et al.
Published: (2024)
by: Lobo, Elita, et al.
Published: (2024)
Manipulating Large Language Models to Increase Product Visibility
by: Kumar, Aounon, et al.
Published: (2024)
by: Kumar, Aounon, et al.
Published: (2024)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
by: Kroeger, Nicholas, et al.
Published: (2023)
by: Kroeger, Nicholas, et al.
Published: (2023)
On the Trade-offs between Adversarial Robustness and Actionable Explanations
by: Krishna, Satyapriya, et al.
Published: (2023)
by: Krishna, Satyapriya, et al.
Published: (2023)
Confronting LLMs with Traditional ML: Rethinking the Fairness of Large Language Models in Tabular Classifications
by: Liu, Yanchen, et al.
Published: (2023)
by: Liu, Yanchen, et al.
Published: (2023)
Towards Uncovering How Large Language Model Works: An Explainability Perspective
by: Zhao, Haiyan, et al.
Published: (2024)
by: Zhao, Haiyan, et al.
Published: (2024)
MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models
by: Han, Tessa, et al.
Published: (2024)
by: Han, Tessa, et al.
Published: (2024)
Certifying LLM Safety against Adversarial Prompting
by: Kumar, Aounon, et al.
Published: (2023)
by: Kumar, Aounon, et al.
Published: (2023)
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach
by: Wojciechowski, Adam, et al.
Published: (2024)
by: Wojciechowski, Adam, et al.
Published: (2024)
Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations
by: Banerjee, Somnath, et al.
Published: (2026)
by: Banerjee, Somnath, et al.
Published: (2026)
Quantifying Generalization Complexity for Large Language Models
by: Qi, Zhenting, et al.
Published: (2024)
by: Qi, Zhenting, et al.
Published: (2024)
FaithLM: Towards Faithful Explanations for Large Language Models
by: Chuang, Yu-Neng, et al.
Published: (2024)
by: Chuang, Yu-Neng, et al.
Published: (2024)
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
by: Li, Aaron J., et al.
Published: (2024)
by: Li, Aaron J., et al.
Published: (2024)
Investigating Training and Generalization in Faithful Self-Explanations of Large Language Models
by: Doi, Tomoki, et al.
Published: (2025)
by: Doi, Tomoki, et al.
Published: (2025)
Self-Improving Language Models with Bidirectional Evolutionary Search
by: Xu, Guowei, et al.
Published: (2026)
by: Xu, Guowei, et al.
Published: (2026)
Evaluating the Reliability of Self-Explanations in Large Language Models
by: Randl, Korbinian, et al.
Published: (2024)
by: Randl, Korbinian, et al.
Published: (2024)
Interpretability Needs a New Paradigm
by: Madsen, Andreas, et al.
Published: (2024)
by: Madsen, Andreas, et al.
Published: (2024)
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
by: Matton, Katie, et al.
Published: (2025)
by: Matton, Katie, et al.
Published: (2025)
Faithful-Patchscopes: Understanding and Mitigating Model Bias in Hidden Representations Explanation of Large Language Models
by: Gong, Xilin, et al.
Published: (2026)
by: Gong, Xilin, et al.
Published: (2026)
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
by: Siegel, Noah Y., et al.
Published: (2024)
by: Siegel, Noah Y., et al.
Published: (2024)
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models
by: Xiong, Zidi, et al.
Published: (2025)
by: Xiong, Zidi, et al.
Published: (2025)
Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language Models
by: Yeo, Wei Jie, et al.
Published: (2024)
by: Yeo, Wei Jie, et al.
Published: (2024)
Large Language Models for Psycholinguistic Plausibility Pretesting
by: Amouyal, Samuel Joseph, et al.
Published: (2024)
by: Amouyal, Samuel Joseph, et al.
Published: (2024)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
by: Li, Aaron J., et al.
Published: (2025)
by: Li, Aaron J., et al.
Published: (2025)
Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability
by: Bhalla, Usha, et al.
Published: (2025)
by: Bhalla, Usha, et al.
Published: (2025)
Self-Critique and Refinement for Faithful Natural Language Explanations
by: Wang, Yingming, et al.
Published: (2025)
by: Wang, Yingming, et al.
Published: (2025)
Polarity-Aware Probing for Quantifying Latent Alignment in Language Models
by: Sadiekh, Sabrina, et al.
Published: (2025)
by: Sadiekh, Sabrina, et al.
Published: (2025)
A Survey of Multilingual Reasoning in Language Models
by: Ghosh, Akash, et al.
Published: (2025)
by: Ghosh, Akash, et al.
Published: (2025)
Illocutionary Explanation Planning for Source-Faithful Explanations in Retrieval-Augmented Language Models
by: Sovrano, Francesco, et al.
Published: (2026)
by: Sovrano, Francesco, et al.
Published: (2026)
Learning Recourse Costs from Pairwise Feature Comparisons
by: Rawal, Kaivalya, et al.
Published: (2024)
by: Rawal, Kaivalya, et al.
Published: (2024)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
by: Qi, Zhenting, et al.
Published: (2024)
by: Qi, Zhenting, et al.
Published: (2024)
Towards Faithful Model Explanation in NLP: A Survey
by: Lyu, Qing, et al.
Published: (2022)
by: Lyu, Qing, et al.
Published: (2022)
From Critique to Clarity: A Pathway to Faithful and Personalized Code Explanations with Large Language Models
by: Xu, Zexing, et al.
Published: (2024)
by: Xu, Zexing, et al.
Published: (2024)
Comparing Explanation Faithfulness between Multilingual and Monolingual Fine-tuned Language Models
by: Zhao, Zhixue, et al.
Published: (2024)
by: Zhao, Zhixue, et al.
Published: (2024)
ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability
by: Sun, Chung-En, et al.
Published: (2025)
by: Sun, Chung-En, et al.
Published: (2025)
EvoLM: In Search of Lost Language Model Training Dynamics
by: Qi, Zhenting, et al.
Published: (2025)
by: Qi, Zhenting, et al.
Published: (2025)
Regularization, Semi-supervision, and Supervision for a Plausible Attention-Based Explanation
by: Nguyen, Duc Hau, et al.
Published: (2025)
by: Nguyen, Duc Hau, et al.
Published: (2025)
Log Probabilities Are a Reliable Estimate of Semantic Plausibility in Base and Instruction-Tuned Language Models
by: Kauf, Carina, et al.
Published: (2024)
by: Kauf, Carina, et al.
Published: (2024)
Similar Items
-
On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models
by: Tanneru, Sree Harsha, et al.
Published: (2024) -
Understanding the Effects of Iterative Prompting on Truthfulness
by: Krishna, Satyapriya, et al.
Published: (2024) -
On the Impact of Fine-Tuning on Chain-of-Thought Reasoning
by: Lobo, Elita, et al.
Published: (2024) -
Manipulating Large Language Models to Increase Product Visibility
by: Kumar, Aounon, et al.
Published: (2024) -
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
by: Kroeger, Nicholas, et al.
Published: (2023)