A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Lingjun, Daumé III, Hal |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Successfully Guiding Humans with Imperfect Instructions by Highlighting Potential Errors and Suggesting Corrections
von: Zhao, Lingjun, et al.
Veröffentlicht: (2024)
von: Zhao, Lingjun, et al.
Veröffentlicht: (2024)
SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models
von: Nghiem, Huy, et al.
Veröffentlicht: (2025)
von: Nghiem, Huy, et al.
Veröffentlicht: (2025)
Pragmatics Meets Culture: Culturally-adapted Artwork Description Generation and Evaluation
von: Zhao, Lingjun, et al.
Veröffentlicht: (2026)
von: Zhao, Lingjun, et al.
Veröffentlicht: (2026)
HateCOT: An Explanation-Enhanced Dataset for Generalizable Offensive Speech Detection via Large Language Models
von: Nghiem, Huy, et al.
Veröffentlicht: (2024)
von: Nghiem, Huy, et al.
Veröffentlicht: (2024)
Can Hallucination Correction Improve Video-Language Alignment?
von: Zhao, Lingjun, et al.
Veröffentlicht: (2025)
von: Zhao, Lingjun, et al.
Veröffentlicht: (2025)
Steering Safely or Off a Cliff? Rethinking Specificity and Robustness in Inference-Time Interventions
von: Goyal, Navita, et al.
Veröffentlicht: (2026)
von: Goyal, Navita, et al.
Veröffentlicht: (2026)
Language Models Predict Empathy Gaps Between Social In-groups and Out-groups
von: Hou, Yu, et al.
Veröffentlicht: (2025)
von: Hou, Yu, et al.
Veröffentlicht: (2025)
"You Gotta be a Doctor, Lin": An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations
von: Nghiem, Huy, et al.
Veröffentlicht: (2024)
von: Nghiem, Huy, et al.
Veröffentlicht: (2024)
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
von: Gor, Maharshi, et al.
Veröffentlicht: (2024)
von: Gor, Maharshi, et al.
Veröffentlicht: (2024)
Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators
von: Cao, Yang Trista, et al.
Veröffentlicht: (2023)
von: Cao, Yang Trista, et al.
Veröffentlicht: (2023)
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
von: Siegel, Noah Y., et al.
Veröffentlicht: (2024)
von: Siegel, Noah Y., et al.
Veröffentlicht: (2024)
When Stereotypes GTG: The Impact of Predictive Text Suggestions on Gender Bias in Human-AI Co-Writing
von: Baumler, Connor, et al.
Veröffentlicht: (2024)
von: Baumler, Connor, et al.
Veröffentlicht: (2024)
Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring
von: Nghiem, Huy, et al.
Veröffentlicht: (2026)
von: Nghiem, Huy, et al.
Veröffentlicht: (2026)
Reheat Nachos for Dinner? Evaluating AI Support for Cross-Cultural Communication of Neologisms
von: Ki, Dayeon, et al.
Veröffentlicht: (2026)
von: Ki, Dayeon, et al.
Veröffentlicht: (2026)
The Impact of Explanations on Fairness in Human-AI Decision-Making: Protected vs Proxy Features
von: Goyal, Navita, et al.
Veröffentlicht: (2023)
von: Goyal, Navita, et al.
Veröffentlicht: (2023)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
von: Nghiem, Huy, et al.
Veröffentlicht: (2025)
von: Nghiem, Huy, et al.
Veröffentlicht: (2025)
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
von: Matton, Katie, et al.
Veröffentlicht: (2025)
von: Matton, Katie, et al.
Veröffentlicht: (2025)
Local Explanations and Self-Explanations for Assessing Faithfulness in black-box LLMs
von: Fragkathoulas, Christos, et al.
Veröffentlicht: (2024)
von: Fragkathoulas, Christos, et al.
Veröffentlicht: (2024)
Comparing Explanation Faithfulness between Multilingual and Monolingual Fine-tuned Language Models
von: Zhao, Zhixue, et al.
Veröffentlicht: (2024)
von: Zhao, Zhixue, et al.
Veröffentlicht: (2024)
Faithfulness and the Notion of Adversarial Sensitivity in NLP Explanations
von: Manna, Supriya, et al.
Veröffentlicht: (2024)
von: Manna, Supriya, et al.
Veröffentlicht: (2024)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
von: Alon, Bar, et al.
Veröffentlicht: (2026)
von: Alon, Bar, et al.
Veröffentlicht: (2026)
Improving Faithfulness of Large Language Models in Summarization via Sliding Generation and Self-Consistency
von: Li, Taiji, et al.
Veröffentlicht: (2024)
von: Li, Taiji, et al.
Veröffentlicht: (2024)
FaithLM: Towards Faithful Explanations for Large Language Models
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
von: Chuang, Yu-Neng, et al.
Veröffentlicht: (2024)
Call Me When Necessary: LLMs can Efficiently and Faithfully Reason over Structured Environments
von: Cheng, Sitao, et al.
Veröffentlicht: (2024)
von: Cheng, Sitao, et al.
Veröffentlicht: (2024)
Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations
von: Quan, Xin, et al.
Veröffentlicht: (2025)
von: Quan, Xin, et al.
Veröffentlicht: (2025)
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach
von: Wojciechowski, Adam, et al.
Veröffentlicht: (2024)
von: Wojciechowski, Adam, et al.
Veröffentlicht: (2024)
Illocutionary Explanation Planning for Source-Faithful Explanations in Retrieval-Augmented Language Models
von: Sovrano, Francesco, et al.
Veröffentlicht: (2026)
von: Sovrano, Francesco, et al.
Veröffentlicht: (2026)
CAuSE: Decoding Multimodal Classifiers using Faithful Natural Language Explanation
von: Bandyopadhyay, Dibyanayan, et al.
Veröffentlicht: (2025)
von: Bandyopadhyay, Dibyanayan, et al.
Veröffentlicht: (2025)
Incorporating Attribution Importance for Improving Faithfulness Metrics
von: Zhao, Zhixue, et al.
Veröffentlicht: (2023)
von: Zhao, Zhixue, et al.
Veröffentlicht: (2023)
Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning
von: Chen, Yanda, et al.
Veröffentlicht: (2024)
von: Chen, Yanda, et al.
Veröffentlicht: (2024)
Towards Faithful Knowledge Graph Explanation Through Deep Alignment in Commonsense Question Answering
von: Zhai, Weihe, et al.
Veröffentlicht: (2023)
von: Zhai, Weihe, et al.
Veröffentlicht: (2023)
ASL STEM Wiki: Dataset and Benchmark for Interpreting STEM Articles
von: Yin, Kayo, et al.
Veröffentlicht: (2024)
von: Yin, Kayo, et al.
Veröffentlicht: (2024)
VinaBench: Benchmark for Faithful and Consistent Visual Narratives
von: Gao, Silin, et al.
Veröffentlicht: (2025)
von: Gao, Silin, et al.
Veröffentlicht: (2025)
Generating Faithful Text From a Knowledge Graph with Noisy Reference Text
von: Hashem, Tahsina, et al.
Veröffentlicht: (2023)
von: Hashem, Tahsina, et al.
Veröffentlicht: (2023)
GeoFaith: A Spatio-Temporal Dual View of Faithful Chain-of-Thought
von: Lv, Weijiang, et al.
Veröffentlicht: (2026)
von: Lv, Weijiang, et al.
Veröffentlicht: (2026)
FaithLens: Detecting and Explaining Faithfulness Hallucination
von: Si, Shuzheng, et al.
Veröffentlicht: (2025)
von: Si, Shuzheng, et al.
Veröffentlicht: (2025)
Learning from Sufficient Rationales: Analysing the Relationship Between Explanation Faithfulness and Token-level Regularisation Strategies
von: Kamp, Jonathan, et al.
Veröffentlicht: (2025)
von: Kamp, Jonathan, et al.
Veröffentlicht: (2025)
Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
von: Siegel, Noah Y., et al.
Veröffentlicht: (2025)
von: Siegel, Noah Y., et al.
Veröffentlicht: (2025)
MAMM-Refine: A Recipe for Improving Faithfulness in Generation with Multi-Agent Collaboration
von: Wan, David, et al.
Veröffentlicht: (2025)
von: Wan, David, et al.
Veröffentlicht: (2025)
PR-CAD: Progressive Refinement for Unified Controllable and Faithful Text-to-CAD Generation with Large Language Models
von: An, Jiyuan, et al.
Veröffentlicht: (2026)
von: An, Jiyuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Successfully Guiding Humans with Imperfect Instructions by Highlighting Potential Errors and Suggesting Corrections
von: Zhao, Lingjun, et al.
Veröffentlicht: (2024) -
SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language Models
von: Nghiem, Huy, et al.
Veröffentlicht: (2025) -
Pragmatics Meets Culture: Culturally-adapted Artwork Description Generation and Evaluation
von: Zhao, Lingjun, et al.
Veröffentlicht: (2026) -
HateCOT: An Explanation-Enhanced Dataset for Generalizable Offensive Speech Detection via Large Language Models
von: Nghiem, Huy, et al.
Veröffentlicht: (2024) -
Can Hallucination Correction Improve Video-Language Alignment?
von: Zhao, Lingjun, et al.
Veröffentlicht: (2025)