NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Bhan, Milan, Vittaut, Jean-Noel, Chesneau, Nicolas, Chandar, Sarath, Lesot, Marie-Jeanne |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Self-AMPLIFY: Improving Small Language Models with Self Post Hoc Explanations
par: Bhan, Milan, et autres
Publié: (2024)
par: Bhan, Milan, et autres
Publié: (2024)
Towards Achieving Concept Completeness for Textual Concept Bottleneck Models
par: Bhan, Milan, et autres
Publié: (2025)
par: Bhan, Milan, et autres
Publié: (2025)
Mitigating Text Toxicity with Counterfactual Generation
par: Bhan, Milan, et autres
Publié: (2024)
par: Bhan, Milan, et autres
Publié: (2024)
Faithfulness Measurable Masked Language Models
par: Madsen, Andreas, et autres
Publié: (2023)
par: Madsen, Andreas, et autres
Publié: (2023)
Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
par: Siegel, Noah Y., et autres
Publié: (2025)
par: Siegel, Noah Y., et autres
Publié: (2025)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
par: Alon, Bar, et autres
Publié: (2026)
par: Alon, Bar, et autres
Publié: (2026)
Evaluating Human Alignment and Model Faithfulness of LLM Rationale
par: Fayyaz, Mohsen, et autres
Publié: (2024)
par: Fayyaz, Mohsen, et autres
Publié: (2024)
Self-Critique and Refinement for Faithful Natural Language Explanations
par: Wang, Yingming, et autres
Publié: (2025)
par: Wang, Yingming, et autres
Publié: (2025)
Local Explanations and Self-Explanations for Assessing Faithfulness in black-box LLMs
par: Fragkathoulas, Christos, et autres
Publié: (2024)
par: Fragkathoulas, Christos, et autres
Publié: (2024)
FaithLM: Towards Faithful Explanations for Large Language Models
par: Chuang, Yu-Neng, et autres
Publié: (2024)
par: Chuang, Yu-Neng, et autres
Publié: (2024)
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
par: Siegel, Noah Y., et autres
Publié: (2024)
par: Siegel, Noah Y., et autres
Publié: (2024)
Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations
par: Quan, Xin, et autres
Publié: (2025)
par: Quan, Xin, et autres
Publié: (2025)
Investigating Training and Generalization in Faithful Self-Explanations of Large Language Models
par: Doi, Tomoki, et autres
Publié: (2025)
par: Doi, Tomoki, et autres
Publié: (2025)
Multilingual Self-Taught Faithfulness Evaluators
par: Alfano, Carlo, et autres
Publié: (2025)
par: Alfano, Carlo, et autres
Publié: (2025)
Can LLMs Produce Faithful Explanations For Fact-checking? Towards Faithful Explainable Fact-Checking via Multi-Agent Debate
par: Kim, Kyungha, et autres
Publié: (2024)
par: Kim, Kyungha, et autres
Publié: (2024)
Towards Faithful Model Explanation in NLP: A Survey
par: Lyu, Qing, et autres
Publié: (2022)
par: Lyu, Qing, et autres
Publié: (2022)
Alignment Reduces Expressed but Not Encoded Gender Bias: A Unified Framework and Study
par: Bouchouchi, Nour, et autres
Publié: (2026)
par: Bouchouchi, Nour, et autres
Publié: (2026)
Faithfulness and the Notion of Adversarial Sensitivity in NLP Explanations
par: Manna, Supriya, et autres
Publié: (2024)
par: Manna, Supriya, et autres
Publié: (2024)
Towards Faithful Knowledge Graph Explanation Through Deep Alignment in Commonsense Question Answering
par: Zhai, Weihe, et autres
Publié: (2023)
par: Zhai, Weihe, et autres
Publié: (2023)
Learning from Self Critique and Refinement for Faithful LLM Summarization
par: Hu, Ting-Yao, et autres
Publié: (2025)
par: Hu, Ting-Yao, et autres
Publié: (2025)
Towards Faithful Explanations for Text Classification with Robustness Improvement and Explanation Guided Training
par: Li, Dongfang, et autres
Publié: (2023)
par: Li, Dongfang, et autres
Publié: (2023)
FaithRL: Learning to Reason Faithfully through Step-Level Faithfulness Maximization
par: Gui, Runquan, et autres
Publié: (2026)
par: Gui, Runquan, et autres
Publié: (2026)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
par: Gur-Arieh, Yoav, et autres
Publié: (2026)
par: Gur-Arieh, Yoav, et autres
Publié: (2026)
Faithful-Patchscopes: Understanding and Mitigating Model Bias in Hidden Representations Explanation of Large Language Models
par: Gong, Xilin, et autres
Publié: (2026)
par: Gong, Xilin, et autres
Publié: (2026)
Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning
par: Jia, Jinghan, et autres
Publié: (2026)
par: Jia, Jinghan, et autres
Publié: (2026)
On Measuring Faithfulness or Self-consistency of Natural Language Explanations
par: Parcalabescu, Letitia, et autres
Publié: (2023)
par: Parcalabescu, Letitia, et autres
Publié: (2023)
Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer
par: Shao, Shun, et autres
Publié: (2026)
par: Shao, Shun, et autres
Publié: (2026)
Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations
par: Banerjee, Somnath, et autres
Publié: (2026)
par: Banerjee, Somnath, et autres
Publié: (2026)
FaithLens: Detecting and Explaining Faithfulness Hallucination
par: Si, Shuzheng, et autres
Publié: (2025)
par: Si, Shuzheng, et autres
Publié: (2025)
Are LLM Decisions Faithful to Verbal Confidence?
par: Wang, Jiawei, et autres
Publié: (2026)
par: Wang, Jiawei, et autres
Publié: (2026)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
par: Mittal, Avni, et autres
Publié: (2026)
par: Mittal, Avni, et autres
Publié: (2026)
Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
par: Agarwal, Chirag, et autres
Publié: (2024)
par: Agarwal, Chirag, et autres
Publié: (2024)
Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
par: Mo, Kaijie, et autres
Publié: (2026)
par: Mo, Kaijie, et autres
Publié: (2026)
Illocutionary Explanation Planning for Source-Faithful Explanations in Retrieval-Augmented Language Models
par: Sovrano, Francesco, et autres
Publié: (2026)
par: Sovrano, Francesco, et autres
Publié: (2026)
FaithCAMERA: Construction of a Faithful Dataset for Ad Text Generation
par: Kato, Akihiko, et autres
Publié: (2024)
par: Kato, Akihiko, et autres
Publié: (2024)
Faithful Model Evaluation for Model-Based Metrics
par: Goyal, Palash, et autres
Publié: (2023)
par: Goyal, Palash, et autres
Publié: (2023)
Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
par: Tian, Changyuan, et autres
Publié: (2026)
par: Tian, Changyuan, et autres
Publié: (2026)
STORYSUMM: Evaluating Faithfulness in Story Summarization
par: Subbiah, Melanie, et autres
Publié: (2024)
par: Subbiah, Melanie, et autres
Publié: (2024)
Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing
par: Yuan, Wenhao, et autres
Publié: (2026)
par: Yuan, Wenhao, et autres
Publié: (2026)
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
par: Matton, Katie, et autres
Publié: (2025)
par: Matton, Katie, et autres
Publié: (2025)
Documents similaires
-
Self-AMPLIFY: Improving Small Language Models with Self Post Hoc Explanations
par: Bhan, Milan, et autres
Publié: (2024) -
Towards Achieving Concept Completeness for Textual Concept Bottleneck Models
par: Bhan, Milan, et autres
Publié: (2025) -
Mitigating Text Toxicity with Counterfactual Generation
par: Bhan, Milan, et autres
Publié: (2024) -
Faithfulness Measurable Masked Language Models
par: Madsen, Andreas, et autres
Publié: (2023) -
Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
par: Siegel, Noah Y., et autres
Publié: (2025)