The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Siegel, Noah Y., Camburu, Oana-Maria, Heess, Nicolas, Perez-Ortiz, Maria |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
von: Siegel, Noah Y., et al.
Veröffentlicht: (2025)
von: Siegel, Noah Y., et al.
Veröffentlicht: (2025)
Identifying Linear Relational Concepts in Large Language Models
von: Chanin, David, et al.
Veröffentlicht: (2023)
von: Chanin, David, et al.
Veröffentlicht: (2023)
FaithLM: Towards Faithful Explanations for Large Language Models
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
SPARSEFIT: Few-shot Prompting with Sparse Fine-tuning for Jointly Generating Predictions and Natural Language Explanations
von: Solano, Jesus, et al.
Veröffentlicht: (2023)
von: Solano, Jesus, et al.
Veröffentlicht: (2023)
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
von: Matton, Katie, et al.
Veröffentlicht: (2025)
von: Matton, Katie, et al.
Veröffentlicht: (2025)
A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanations
von: Zhao, Lingjun, et al.
Veröffentlicht: (2025)
von: Zhao, Lingjun, et al.
Veröffentlicht: (2025)
Illocutionary Explanation Planning for Source-Faithful Explanations in Retrieval-Augmented Language Models
von: Sovrano, Francesco, et al.
Veröffentlicht: (2026)
von: Sovrano, Francesco, et al.
Veröffentlicht: (2026)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
von: Alon, Bar, et al.
Veröffentlicht: (2026)
von: Alon, Bar, et al.
Veröffentlicht: (2026)
Comparing Explanation Faithfulness between Multilingual and Monolingual Fine-tuned Language Models
von: Zhao, Zhixue, et al.
Veröffentlicht: (2024)
von: Zhao, Zhixue, et al.
Veröffentlicht: (2024)
Faithfulness and the Notion of Adversarial Sensitivity in NLP Explanations
von: Manna, Supriya, et al.
Veröffentlicht: (2024)
von: Manna, Supriya, et al.
Veröffentlicht: (2024)
Closing the Confidence-Faithfulness Gap in Large Language Models
von: Miao, Miranda Muqing, et al.
Veröffentlicht: (2026)
von: Miao, Miranda Muqing, et al.
Veröffentlicht: (2026)
FaithUn: Toward Faithful Forgetting in Language Models by Investigating the Interconnectedness of Knowledge
von: Yang, Nakyeong, et al.
Veröffentlicht: (2025)
von: Yang, Nakyeong, et al.
Veröffentlicht: (2025)
Local Explanations and Self-Explanations for Assessing Faithfulness in black-box LLMs
von: Fragkathoulas, Christos, et al.
Veröffentlicht: (2024)
von: Fragkathoulas, Christos, et al.
Veröffentlicht: (2024)
From Critique to Clarity: A Pathway to Faithful and Personalized Code Explanations with Large Language Models
von: Xu, Zexing, et al.
Veröffentlicht: (2024)
von: Xu, Zexing, et al.
Veröffentlicht: (2024)
A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
von: Mayne, Harry, et al.
Veröffentlicht: (2026)
von: Mayne, Harry, et al.
Veröffentlicht: (2026)
CAuSE: Decoding Multimodal Classifiers using Faithful Natural Language Explanation
von: Bandyopadhyay, Dibyanayan, et al.
Veröffentlicht: (2025)
von: Bandyopadhyay, Dibyanayan, et al.
Veröffentlicht: (2025)
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach
von: Wojciechowski, Adam, et al.
Veröffentlicht: (2024)
von: Wojciechowski, Adam, et al.
Veröffentlicht: (2024)
Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning
von: Luo, Linhao, et al.
Veröffentlicht: (2023)
von: Luo, Linhao, et al.
Veröffentlicht: (2023)
Transformer Circuit Faithfulness Metrics are not Robust
von: Miller, Joseph, et al.
Veröffentlicht: (2024)
von: Miller, Joseph, et al.
Veröffentlicht: (2024)
PR-CAD: Progressive Refinement for Unified Controllable and Faithful Text-to-CAD Generation with Large Language Models
von: An, Jiyuan, et al.
Veröffentlicht: (2026)
von: An, Jiyuan, et al.
Veröffentlicht: (2026)
Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations
von: Quan, Xin, et al.
Veröffentlicht: (2025)
von: Quan, Xin, et al.
Veröffentlicht: (2025)
FaithLens: Detecting and Explaining Faithfulness Hallucination
von: Si, Shuzheng, et al.
Veröffentlicht: (2025)
von: Si, Shuzheng, et al.
Veröffentlicht: (2025)
Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language Models
von: Yang, Chenxu, et al.
Veröffentlicht: (2025)
von: Yang, Chenxu, et al.
Veröffentlicht: (2025)
SPD-Faith Bench: Diagnosing and Improving Faithfulness in Chain-of-Thought for Multimodal Large Language Models
von: Lv, Weijiang, et al.
Veröffentlicht: (2026)
von: Lv, Weijiang, et al.
Veröffentlicht: (2026)
Incorporating Attribution Importance for Improving Faithfulness Metrics
von: Zhao, Zhixue, et al.
Veröffentlicht: (2023)
von: Zhao, Zhixue, et al.
Veröffentlicht: (2023)
Prompt-Response Semantic Divergence Metrics for Faithfulness Hallucination and Misalignment Detection in Large Language Models
von: Halperin, Igor
Veröffentlicht: (2025)
von: Halperin, Igor
Veröffentlicht: (2025)
To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External Contexts
von: Huang, Yukun, et al.
Veröffentlicht: (2024)
von: Huang, Yukun, et al.
Veröffentlicht: (2024)
A Causal Lens for Evaluating Faithfulness Metrics
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
Faithfulness-QA: A Counterfactual Entity Substitution Dataset for Training Context-Faithful RAG Models
von: Ju, Li, et al.
Veröffentlicht: (2026)
von: Ju, Li, et al.
Veröffentlicht: (2026)
Improving Faithfulness of Large Language Models in Summarization via Sliding Generation and Self-Consistency
von: Li, Taiji, et al.
Veröffentlicht: (2024)
von: Li, Taiji, et al.
Veröffentlicht: (2024)
Investigating Context-Faithfulness in Large Language Models: The Roles of Memory Strength and Evidence Style
von: Li, Yuepei, et al.
Veröffentlicht: (2024)
von: Li, Yuepei, et al.
Veröffentlicht: (2024)
FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
von: Ming, Yifei, et al.
Veröffentlicht: (2024)
von: Ming, Yifei, et al.
Veröffentlicht: (2024)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
von: Anand, Nikhil, et al.
Veröffentlicht: (2026)
von: Anand, Nikhil, et al.
Veröffentlicht: (2026)
TrueBrief: Faithful Summarization through Small Language Models
von: Lakara, Kumud, et al.
Veröffentlicht: (2025)
von: Lakara, Kumud, et al.
Veröffentlicht: (2025)
GeoFaith: A Spatio-Temporal Dual View of Faithful Chain-of-Thought
von: Lv, Weijiang, et al.
Veröffentlicht: (2026)
von: Lv, Weijiang, et al.
Veröffentlicht: (2026)
Towards Faithful Knowledge Graph Explanation Through Deep Alignment in Commonsense Question Answering
von: Zhai, Weihe, et al.
Veröffentlicht: (2023)
von: Zhai, Weihe, et al.
Veröffentlicht: (2023)
DeepFaith: A Domain-Free and Model-Agnostic Unified Framework for Highly Faithful Explanations
von: Guo, Yuhan, et al.
Veröffentlicht: (2025)
von: Guo, Yuhan, et al.
Veröffentlicht: (2025)
Generating Faithful Text From a Knowledge Graph with Noisy Reference Text
von: Hashem, Tahsina, et al.
Veröffentlicht: (2023)
von: Hashem, Tahsina, et al.
Veröffentlicht: (2023)
FiDeLiS: Faithful Reasoning in Large Language Model for Knowledge Graph Question Answering
von: Sui, Yuan, et al.
Veröffentlicht: (2024)
von: Sui, Yuan, et al.
Veröffentlicht: (2024)
Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement Learning
von: Si, Shuzheng, et al.
Veröffentlicht: (2025)
von: Si, Shuzheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
von: Siegel, Noah Y., et al.
Veröffentlicht: (2025) -
Identifying Linear Relational Concepts in Large Language Models
von: Chanin, David, et al.
Veröffentlicht: (2023) -
FaithLM: Towards Faithful Explanations for Large Language Models
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024) -
SPARSEFIT: Few-shot Prompting with Sparse Fine-tuning for Jointly Generating Predictions and Natural Language Explanations
von: Solano, Jesus, et al.
Veröffentlicht: (2023) -
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
von: Matton, Katie, et al.
Veröffentlicht: (2025)