Local Explanations and Self-Explanations for Assessing Faithfulness in black-box LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Fragkathoulas, Christos, Chlapanis, Odysseas S. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans
di: Chlapanis, Odysseas S., et al.
Pubblicazione: (2026)
di: Chlapanis, Odysseas S., et al.
Pubblicazione: (2026)
LAR-ECHR: A New Legal Argument Reasoning Task and Dataset for Cases of the European Court of Human Rights
di: Chlapanis, Odysseas S., et al.
Pubblicazione: (2024)
di: Chlapanis, Odysseas S., et al.
Pubblicazione: (2024)
Faithfulness and the Notion of Adversarial Sensitivity in NLP Explanations
di: Manna, Supriya, et al.
Pubblicazione: (2024)
di: Manna, Supriya, et al.
Pubblicazione: (2024)
Illocutionary Explanation Planning for Source-Faithful Explanations in Retrieval-Augmented Language Models
di: Sovrano, Francesco, et al.
Pubblicazione: (2026)
di: Sovrano, Francesco, et al.
Pubblicazione: (2026)
FaithLM: Towards Faithful Explanations for Large Language Models
di: Chuang, Yu-Neng, et al.
Pubblicazione: (2024)
di: Chuang, Yu-Neng, et al.
Pubblicazione: (2024)
Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
di: Siegel, Noah Y., et al.
Pubblicazione: (2025)
di: Siegel, Noah Y., et al.
Pubblicazione: (2025)
Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations
di: Quan, Xin, et al.
Pubblicazione: (2025)
di: Quan, Xin, et al.
Pubblicazione: (2025)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
di: Alon, Bar, et al.
Pubblicazione: (2026)
di: Alon, Bar, et al.
Pubblicazione: (2026)
Distilling Text Style Transfer With Self-Explanation From LLMs
di: Zhang, Chiyu, et al.
Pubblicazione: (2024)
di: Zhang, Chiyu, et al.
Pubblicazione: (2024)
FACEGroup: Feasible and Actionable Counterfactual Explanations for Group Fairness
di: Fragkathoulas, Christos, et al.
Pubblicazione: (2024)
di: Fragkathoulas, Christos, et al.
Pubblicazione: (2024)
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
di: Siegel, Noah Y., et al.
Pubblicazione: (2024)
di: Siegel, Noah Y., et al.
Pubblicazione: (2024)
Comparing Explanation Faithfulness between Multilingual and Monolingual Fine-tuned Language Models
di: Zhao, Zhixue, et al.
Pubblicazione: (2024)
di: Zhao, Zhixue, et al.
Pubblicazione: (2024)
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach
di: Wojciechowski, Adam, et al.
Pubblicazione: (2024)
di: Wojciechowski, Adam, et al.
Pubblicazione: (2024)
CAuSE: Decoding Multimodal Classifiers using Faithful Natural Language Explanation
di: Bandyopadhyay, Dibyanayan, et al.
Pubblicazione: (2025)
di: Bandyopadhyay, Dibyanayan, et al.
Pubblicazione: (2025)
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
di: Matton, Katie, et al.
Pubblicazione: (2025)
di: Matton, Katie, et al.
Pubblicazione: (2025)
A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanations
di: Zhao, Lingjun, et al.
Pubblicazione: (2025)
di: Zhao, Lingjun, et al.
Pubblicazione: (2025)
Towards Faithful Knowledge Graph Explanation Through Deep Alignment in Commonsense Question Answering
di: Zhai, Weihe, et al.
Pubblicazione: (2023)
di: Zhai, Weihe, et al.
Pubblicazione: (2023)
Anchored Alignment for Self-Explanations Enhancement
di: Villa-Arenas, Luis Felipe, et al.
Pubblicazione: (2024)
di: Villa-Arenas, Luis Felipe, et al.
Pubblicazione: (2024)
LLM Self-Explanations Fail Semantic Invariance
di: Szeider, Stefan
Pubblicazione: (2026)
di: Szeider, Stefan
Pubblicazione: (2026)
Self-Explanation in Social AI Agents
di: Basappa, Rhea, et al.
Pubblicazione: (2025)
di: Basappa, Rhea, et al.
Pubblicazione: (2025)
Digital Socrates: Evaluating LLMs through Explanation Critiques
di: Gu, Yuling, et al.
Pubblicazione: (2023)
di: Gu, Yuling, et al.
Pubblicazione: (2023)
On Measuring Faithfulness or Self-consistency of Natural Language Explanations
di: Parcalabescu, Letitia, et al.
Pubblicazione: (2023)
di: Parcalabescu, Letitia, et al.
Pubblicazione: (2023)
Learning from Sufficient Rationales: Analysing the Relationship Between Explanation Faithfulness and Token-level Regularisation Strategies
di: Kamp, Jonathan, et al.
Pubblicazione: (2025)
di: Kamp, Jonathan, et al.
Pubblicazione: (2025)
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
di: Mahale, Ajay Pravin
Pubblicazione: (2026)
di: Mahale, Ajay Pravin
Pubblicazione: (2026)
No Need for Explanations: LLMs can implicitly learn from mistakes in-context
di: Alazraki, Lisa, et al.
Pubblicazione: (2025)
di: Alazraki, Lisa, et al.
Pubblicazione: (2025)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
di: Mayne, Harry, et al.
Pubblicazione: (2025)
di: Mayne, Harry, et al.
Pubblicazione: (2025)
From Critique to Clarity: A Pathway to Faithful and Personalized Code Explanations with Large Language Models
di: Xu, Zexing, et al.
Pubblicazione: (2024)
di: Xu, Zexing, et al.
Pubblicazione: (2024)
PLEX: Perturbation-free Local Explanations for LLM-Based Text Classification
di: Rahulamathavan, Yogachandran, et al.
Pubblicazione: (2025)
di: Rahulamathavan, Yogachandran, et al.
Pubblicazione: (2025)
Assessing Large Language Models for Online Extremism Research: Identification, Explanation, and New Knowledge
di: Dong, Beidi, et al.
Pubblicazione: (2024)
di: Dong, Beidi, et al.
Pubblicazione: (2024)
Explaining Black-box Language Models with Knowledge Probing Systems: A Post-hoc Explanation Perspective
di: Zhao, Yunxiao, et al.
Pubblicazione: (2025)
di: Zhao, Yunxiao, et al.
Pubblicazione: (2025)
XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMs
di: Chen, Zichen, et al.
Pubblicazione: (2023)
di: Chen, Zichen, et al.
Pubblicazione: (2023)
Reasoning with Natural Language Explanations
di: Valentino, Marco, et al.
Pubblicazione: (2024)
di: Valentino, Marco, et al.
Pubblicazione: (2024)
LLMs for XAI: Future Directions for Explaining Explanations
di: Zytek, Alexandra, et al.
Pubblicazione: (2024)
di: Zytek, Alexandra, et al.
Pubblicazione: (2024)
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
di: Toker, Gilat, et al.
Pubblicazione: (2026)
di: Toker, Gilat, et al.
Pubblicazione: (2026)
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
di: Hamman, Faisal, et al.
Pubblicazione: (2025)
di: Hamman, Faisal, et al.
Pubblicazione: (2025)
Take It Easy: Label-Adaptive Self-Rationalization for Fact Verification and Explanation Generation
di: Yang, Jing, et al.
Pubblicazione: (2024)
di: Yang, Jing, et al.
Pubblicazione: (2024)
AUEB-Archimedes at RIRAG-2025: Is obligation concatenation really all you need?
di: Chasandras, Ioannis, et al.
Pubblicazione: (2024)
di: Chasandras, Ioannis, et al.
Pubblicazione: (2024)
Archimedes-AUEB at SemEval-2024 Task 5: LLM explains Civil Procedure
di: Chlapanis, Odysseas S., et al.
Pubblicazione: (2024)
di: Chlapanis, Odysseas S., et al.
Pubblicazione: (2024)
CAVE: Controllable Authorship Verification Explanations
di: Ramnath, Sahana, et al.
Pubblicazione: (2024)
di: Ramnath, Sahana, et al.
Pubblicazione: (2024)
Support-Contra Asymmetry in LLM Explanations
di: Patil, Avinash
Pubblicazione: (2025)
di: Patil, Avinash
Pubblicazione: (2025)
Documenti analoghi
-
The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans
di: Chlapanis, Odysseas S., et al.
Pubblicazione: (2026) -
LAR-ECHR: A New Legal Argument Reasoning Task and Dataset for Cases of the European Court of Human Rights
di: Chlapanis, Odysseas S., et al.
Pubblicazione: (2024) -
Faithfulness and the Notion of Adversarial Sensitivity in NLP Explanations
di: Manna, Supriya, et al.
Pubblicazione: (2024) -
Illocutionary Explanation Planning for Source-Faithful Explanations in Retrieval-Augmented Language Models
di: Sovrano, Francesco, et al.
Pubblicazione: (2026) -
FaithLM: Towards Faithful Explanations for Large Language Models
di: Chuang, Yu-Neng, et al.
Pubblicazione: (2024)