FaithLM: Towards Faithful Explanations for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Chuang, Yu-Neng, Wang, Guanchu, Chang, Chia-Yuan, Tang, Ruixiang, Zhong, Shaochen, Yang, Fan, Du, Mengnan, Cai, Xuanting, Braverman, Vladimir, Hu, Xia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TVE: Learning Meta-attribution for Transferable Vision Explainer
by: Wang, Guanchu, et al.
Published: (2023)
by: Wang, Guanchu, et al.
Published: (2023)
Assessing and Enhancing Large Language Models in Rare Disease Question-answering
by: Wang, Guanchu, et al.
Published: (2024)
by: Wang, Guanchu, et al.
Published: (2024)
Confident or Seek Stronger: Exploring Uncertainty-Based On-device LLM Routing From Benchmarking to Generalization
by: Chuang, Yu-Neng, et al.
Published: (2025)
by: Chuang, Yu-Neng, et al.
Published: (2025)
Self-ensemble: Mitigating Confidence Mis-calibration for Large Language Models
by: Xu, Zicheng, et al.
Published: (2025)
by: Xu, Zicheng, et al.
Published: (2025)
Taylor Unswift: Secured Weight Release for Large Language Models via Taylor Expansion
by: Wang, Guanchu, et al.
Published: (2024)
by: Wang, Guanchu, et al.
Published: (2024)
AutoL2S: Auto Long-Short Reasoning for Efficient Large Language Models
by: Luo, Feng, et al.
Published: (2025)
by: Luo, Feng, et al.
Published: (2025)
On the Faithfulness of Vision Transformer Explanations
by: Wu, Junyi, et al.
Published: (2024)
by: Wu, Junyi, et al.
Published: (2024)
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
by: Luo, Feng, et al.
Published: (2026)
by: Luo, Feng, et al.
Published: (2026)
Towards Faithful Explanations: Boosting Rationalization with Shortcuts Discovery
by: Yue, Linan, et al.
Published: (2024)
by: Yue, Linan, et al.
Published: (2024)
Towards Faithful Explanations for Text Classification with Robustness Improvement and Explanation Guided Training
by: Li, Dongfang, et al.
Published: (2023)
by: Li, Dongfang, et al.
Published: (2023)
Toward Faithful Explanations in Acoustic Anomaly Detection
by: Elrashid, Maab, et al.
Published: (2026)
by: Elrashid, Maab, et al.
Published: (2026)
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
by: Sui, Yang, et al.
Published: (2025)
by: Sui, Yang, et al.
Published: (2025)
DTS: Enhancing Large Reasoning Models via Decoding Tree Sketching
by: Xu, Zicheng, et al.
Published: (2025)
by: Xu, Zicheng, et al.
Published: (2025)
Towards Faithful Model Explanation in NLP: A Survey
by: Lyu, Qing, et al.
Published: (2022)
by: Lyu, Qing, et al.
Published: (2022)
Token Transformation Matters: Towards Faithful Post-hoc Explanation for Vision Transformer
by: Wu, Junyi, et al.
Published: (2024)
by: Wu, Junyi, et al.
Published: (2024)
SAM-Sode: Towards Faithful Explanations for Tiny Bacteria Detection
by: Tan, Wanying, et al.
Published: (2026)
by: Tan, Wanying, et al.
Published: (2026)
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
by: Siegel, Noah Y., et al.
Published: (2024)
by: Siegel, Noah Y., et al.
Published: (2024)
FIRE: Faithful Interpretable Recommendation Explanations
by: Sani, S. M. F., et al.
Published: (2025)
by: Sani, S. M. F., et al.
Published: (2025)
Faithful Counterfactual Visual Explanations (FCVE)
by: Khan, Bismillah, et al.
Published: (2025)
by: Khan, Bismillah, et al.
Published: (2025)
Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation
by: Zhao, Haiyan, et al.
Published: (2026)
by: Zhao, Haiyan, et al.
Published: (2026)
NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
by: Bhan, Milan, et al.
Published: (2025)
by: Bhan, Milan, et al.
Published: (2025)
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
by: Matton, Katie, et al.
Published: (2025)
by: Matton, Katie, et al.
Published: (2025)
Can LLMs Produce Faithful Explanations For Fact-checking? Towards Faithful Explainable Fact-Checking via Multi-Agent Debate
by: Kim, Kyungha, et al.
Published: (2024)
by: Kim, Kyungha, et al.
Published: (2024)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
by: Alon, Bar, et al.
Published: (2026)
by: Alon, Bar, et al.
Published: (2026)
DeepFaith: A Domain-Free and Model-Agnostic Unified Framework for Highly Faithful Explanations
by: Guo, Yuhan, et al.
Published: (2025)
by: Guo, Yuhan, et al.
Published: (2025)
KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches
by: Yuan, Jiayi, et al.
Published: (2024)
by: Yuan, Jiayi, et al.
Published: (2024)
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
by: Liu, Zirui, et al.
Published: (2024)
by: Liu, Zirui, et al.
Published: (2024)
QGShap: Quantum Acceleration for Faithful GNN Explanations
by: Jena, Haribandhu, et al.
Published: (2025)
by: Jena, Haribandhu, et al.
Published: (2025)
"Faithful to What?" On the Limits of Fidelity-Based Explanations
by: Eshbaugh, Jackson
Published: (2025)
by: Eshbaugh, Jackson
Published: (2025)
Evaluating Readability and Faithfulness of Concept-based Explanations
by: Li, Meng, et al.
Published: (2024)
by: Li, Meng, et al.
Published: (2024)
Measuring the (Un)Faithfulness of Concept-Based Explanations
by: Kumar, Shubham, et al.
Published: (2025)
by: Kumar, Shubham, et al.
Published: (2025)
Faithfulness and the Notion of Adversarial Sensitivity in NLP Explanations
by: Manna, Supriya, et al.
Published: (2024)
by: Manna, Supriya, et al.
Published: (2024)
Sparse and Faithful Explanations Without Sparse Models
by: Sun, Yiyang, et al.
Published: (2024)
by: Sun, Yiyang, et al.
Published: (2024)
Faithful Group Shapley Value
by: Lee, Kiljae, et al.
Published: (2025)
by: Lee, Kiljae, et al.
Published: (2025)
Investigating Training and Generalization in Faithful Self-Explanations of Large Language Models
by: Doi, Tomoki, et al.
Published: (2025)
by: Doi, Tomoki, et al.
Published: (2025)
Winner-Take-All Column Row Sampling for Memory Efficient Adaptation of Language Model
by: Liu, Zirui, et al.
Published: (2023)
by: Liu, Zirui, et al.
Published: (2023)
FaithLens: Detecting and Explaining Faithfulness Hallucination
by: Si, Shuzheng, et al.
Published: (2025)
by: Si, Shuzheng, et al.
Published: (2025)
SPD-Faith Bench: Diagnosing and Improving Faithfulness in Chain-of-Thought for Multimodal Large Language Models
by: Lv, Weijiang, et al.
Published: (2026)
by: Lv, Weijiang, et al.
Published: (2026)
Faith in Law, Law in Faith
Published: (2024)
Published: (2024)
Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language Models
by: Yeo, Wei Jie, et al.
Published: (2024)
by: Yeo, Wei Jie, et al.
Published: (2024)
Similar Items
-
TVE: Learning Meta-attribution for Transferable Vision Explainer
by: Wang, Guanchu, et al.
Published: (2023) -
Assessing and Enhancing Large Language Models in Rare Disease Question-answering
by: Wang, Guanchu, et al.
Published: (2024) -
Confident or Seek Stronger: Exploring Uncertainty-Based On-device LLM Routing From Benchmarking to Generalization
by: Chuang, Yu-Neng, et al.
Published: (2025) -
Self-ensemble: Mitigating Confidence Mis-calibration for Large Language Models
by: Xu, Zicheng, et al.
Published: (2025) -
Taylor Unswift: Secured Weight Release for Large Language Models via Taylor Expansion
by: Wang, Guanchu, et al.
Published: (2024)