Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Yeo, Wei Jie, Satapathy, Ranjan, Cambria, Erik |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How Interpretable are Reasoning Explanations from Prompting Large Language Models?
di: Yeo, Wei Jie, et al.
Pubblicazione: (2024)
di: Yeo, Wei Jie, et al.
Pubblicazione: (2024)
Self-training Large Language Models through Knowledge Detection
di: Yeo, Wei Jie, et al.
Pubblicazione: (2024)
di: Yeo, Wei Jie, et al.
Pubblicazione: (2024)
Plausible Extractive Rationalization through Semi-Supervised Entailment Signal
di: Yeo, Wei Jie, et al.
Pubblicazione: (2024)
di: Yeo, Wei Jie, et al.
Pubblicazione: (2024)
Mitigating Jailbreaks with Intent-Aware LLMs
di: Yeo, Wei Jie, et al.
Pubblicazione: (2025)
di: Yeo, Wei Jie, et al.
Pubblicazione: (2025)
Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
di: Prakash, Nirmalendu, et al.
Pubblicazione: (2025)
di: Prakash, Nirmalendu, et al.
Pubblicazione: (2025)
Understanding Refusal in Language Models with Sparse Autoencoders
di: Yeo, Wei Jie, et al.
Pubblicazione: (2025)
di: Yeo, Wei Jie, et al.
Pubblicazione: (2025)
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
di: Yeo, Wei Jie, et al.
Pubblicazione: (2025)
di: Yeo, Wei Jie, et al.
Pubblicazione: (2025)
Explainable Natural Language Processing for Corporate Sustainability Analysis
di: Ong, Keane, et al.
Pubblicazione: (2024)
di: Ong, Keane, et al.
Pubblicazione: (2024)
FaithLM: Towards Faithful Explanations for Large Language Models
di: Chuang, Yu-Neng, et al.
Pubblicazione: (2024)
di: Chuang, Yu-Neng, et al.
Pubblicazione: (2024)
A Systematic Analysis of Biases in Large Language Models
di: Zhang, Xulang, et al.
Pubblicazione: (2025)
di: Zhang, Xulang, et al.
Pubblicazione: (2025)
Self-Critique and Refinement for Faithful Natural Language Explanations
di: Wang, Yingming, et al.
Pubblicazione: (2025)
di: Wang, Yingming, et al.
Pubblicazione: (2025)
Beyond Correlation: Refutation-Validated Aspect-Based Sentiment Analysis for Explainable Energy Market Returns
di: van der Heever, Wihan, et al.
Pubblicazione: (2026)
di: van der Heever, Wihan, et al.
Pubblicazione: (2026)
FinXABSA: Explainable Finance through Aspect-Based Sentiment Analysis
di: Ong, Keane, et al.
Pubblicazione: (2023)
di: Ong, Keane, et al.
Pubblicazione: (2023)
Investigating Training and Generalization in Faithful Self-Explanations of Large Language Models
di: Doi, Tomoki, et al.
Pubblicazione: (2025)
di: Doi, Tomoki, et al.
Pubblicazione: (2025)
Natural Language Counterfactual Explanations for Graphs Using Large Language Models
di: Giorgi, Flavio, et al.
Pubblicazione: (2024)
di: Giorgi, Flavio, et al.
Pubblicazione: (2024)
Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
di: Agarwal, Chirag, et al.
Pubblicazione: (2024)
di: Agarwal, Chirag, et al.
Pubblicazione: (2024)
Large Language Models for Automated Open-domain Scientific Hypotheses Discovery
di: Yang, Zonglin, et al.
Pubblicazione: (2023)
di: Yang, Zonglin, et al.
Pubblicazione: (2023)
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
di: Matton, Katie, et al.
Pubblicazione: (2025)
di: Matton, Katie, et al.
Pubblicazione: (2025)
Large Language Models for Few-Shot Named Entity Recognition
di: Zhao, Yufei, et al.
Pubblicazione: (2018)
di: Zhao, Yufei, et al.
Pubblicazione: (2018)
XAI meets LLMs: A Survey of the Relation between Explainable AI and Large Language Models
di: Cambria, Erik, et al.
Pubblicazione: (2024)
di: Cambria, Erik, et al.
Pubblicazione: (2024)
ESGSenticNet: A Neurosymbolic Knowledge Base for Corporate Sustainability Analysis
di: Ong, Keane, et al.
Pubblicazione: (2025)
di: Ong, Keane, et al.
Pubblicazione: (2025)
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach
di: Wojciechowski, Adam, et al.
Pubblicazione: (2024)
di: Wojciechowski, Adam, et al.
Pubblicazione: (2024)
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
di: Siegel, Noah Y., et al.
Pubblicazione: (2024)
di: Siegel, Noah Y., et al.
Pubblicazione: (2024)
Faithful-Patchscopes: Understanding and Mitigating Model Bias in Hidden Representations Explanation of Large Language Models
di: Gong, Xilin, et al.
Pubblicazione: (2026)
di: Gong, Xilin, et al.
Pubblicazione: (2026)
Logical Reasoning over Natural Language as Knowledge Representation: A Survey
di: Yang, Zonglin, et al.
Pubblicazione: (2023)
di: Yang, Zonglin, et al.
Pubblicazione: (2023)
Towards Faithful Model Explanation in NLP: A Survey
di: Lyu, Qing, et al.
Pubblicazione: (2022)
di: Lyu, Qing, et al.
Pubblicazione: (2022)
CAuSE: Decoding Multimodal Classifiers using Faithful Natural Language Explanation
di: Bandyopadhyay, Dibyanayan, et al.
Pubblicazione: (2025)
di: Bandyopadhyay, Dibyanayan, et al.
Pubblicazione: (2025)
Situated Natural Language Explanations
di: Zhu, Zining, et al.
Pubblicazione: (2023)
di: Zhu, Zining, et al.
Pubblicazione: (2023)
Harnessing Large Language Models for Scientific Novelty Detection
di: Liu, Yan, et al.
Pubblicazione: (2025)
di: Liu, Yan, et al.
Pubblicazione: (2025)
Towards Transparent Reasoning: What Drives Faithfulness in Large Language Models?
di: McMillan, Teague, et al.
Pubblicazione: (2025)
di: McMillan, Teague, et al.
Pubblicazione: (2025)
LExT: Towards Evaluating Trustworthiness of Natural Language Explanations
di: Shailya, Krithi, et al.
Pubblicazione: (2025)
di: Shailya, Krithi, et al.
Pubblicazione: (2025)
On Measuring Faithfulness or Self-consistency of Natural Language Explanations
di: Parcalabescu, Letitia, et al.
Pubblicazione: (2023)
di: Parcalabescu, Letitia, et al.
Pubblicazione: (2023)
Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond
di: Xu, Fangzhi, et al.
Pubblicazione: (2023)
di: Xu, Fangzhi, et al.
Pubblicazione: (2023)
Illocutionary Explanation Planning for Source-Faithful Explanations in Retrieval-Augmented Language Models
di: Sovrano, Francesco, et al.
Pubblicazione: (2026)
di: Sovrano, Francesco, et al.
Pubblicazione: (2026)
A Survey of Large Language Models for Healthcare: from Data, Technology, and Applications to Accountability and Ethics
di: He, Kai, et al.
Pubblicazione: (2023)
di: He, Kai, et al.
Pubblicazione: (2023)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
di: Anand, Nikhil, et al.
Pubblicazione: (2026)
di: Anand, Nikhil, et al.
Pubblicazione: (2026)
From Critique to Clarity: A Pathway to Faithful and Personalized Code Explanations with Large Language Models
di: Xu, Zexing, et al.
Pubblicazione: (2024)
di: Xu, Zexing, et al.
Pubblicazione: (2024)
Using Natural Language Explanations to Rescale Human Judgments
di: Wadhwa, Manya, et al.
Pubblicazione: (2023)
di: Wadhwa, Manya, et al.
Pubblicazione: (2023)
Quantifying Uncertainty in Natural Language Explanations of Large Language Models for Question Answering
di: Li, Yangyi, et al.
Pubblicazione: (2025)
di: Li, Yangyi, et al.
Pubblicazione: (2025)
CELL your Model: Contrastive Explanations for Large Language Models
di: Luss, Ronny, et al.
Pubblicazione: (2024)
di: Luss, Ronny, et al.
Pubblicazione: (2024)
Documenti analoghi
-
How Interpretable are Reasoning Explanations from Prompting Large Language Models?
di: Yeo, Wei Jie, et al.
Pubblicazione: (2024) -
Self-training Large Language Models through Knowledge Detection
di: Yeo, Wei Jie, et al.
Pubblicazione: (2024) -
Plausible Extractive Rationalization through Semi-Supervised Entailment Signal
di: Yeo, Wei Jie, et al.
Pubblicazione: (2024) -
Mitigating Jailbreaks with Intent-Aware LLMs
di: Yeo, Wei Jie, et al.
Pubblicazione: (2025) -
Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
di: Prakash, Nirmalendu, et al.
Pubblicazione: (2025)