Self-Critique and Refinement for Faithful Natural Language Explanations
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Yingming, Atanasova, Pepa |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
por: Sun, Jingyi, et al.
Publicado: (2024)
por: Sun, Jingyi, et al.
Publicado: (2024)
Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
por: Sun, Jingyi, et al.
Publicado: (2025)
por: Sun, Jingyi, et al.
Publicado: (2025)
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
por: Sun, Jingyi, et al.
Publicado: (2026)
por: Sun, Jingyi, et al.
Publicado: (2026)
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
por: Wang, Qianli, et al.
Publicado: (2026)
por: Wang, Qianli, et al.
Publicado: (2026)
Learning from Self Critique and Refinement for Faithful LLM Summarization
por: Hu, Ting-Yao, et al.
Publicado: (2025)
por: Hu, Ting-Yao, et al.
Publicado: (2025)
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods
por: Yu, Haeun, et al.
Publicado: (2024)
por: Yu, Haeun, et al.
Publicado: (2024)
Graph-Guided Textual Explanation Generation Framework
por: Yuan, Shuzhou, et al.
Publicado: (2024)
por: Yuan, Shuzhou, et al.
Publicado: (2024)
Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement
por: Islam, Sekh Mainul, et al.
Publicado: (2025)
por: Islam, Sekh Mainul, et al.
Publicado: (2025)
Dancing with Critiques: Enhancing LLM Reasoning with Stepwise Natural Language Self-Critique
por: Li, Yansi, et al.
Publicado: (2025)
por: Li, Yansi, et al.
Publicado: (2025)
On Measuring Faithfulness or Self-consistency of Natural Language Explanations
por: Parcalabescu, Letitia, et al.
Publicado: (2023)
por: Parcalabescu, Letitia, et al.
Publicado: (2023)
From Critique to Clarity: A Pathway to Faithful and Personalized Code Explanations with Large Language Models
por: Xu, Zexing, et al.
Publicado: (2024)
por: Xu, Zexing, et al.
Publicado: (2024)
Investigating Training and Generalization in Faithful Self-Explanations of Large Language Models
por: Doi, Tomoki, et al.
Publicado: (2025)
por: Doi, Tomoki, et al.
Publicado: (2025)
Training Language Model to Critique for Better Refinement
por: Yu, Tianshu, et al.
Publicado: (2025)
por: Yu, Tianshu, et al.
Publicado: (2025)
NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
por: Bhan, Milan, et al.
Publicado: (2025)
por: Bhan, Milan, et al.
Publicado: (2025)
CAuSE: Decoding Multimodal Classifiers using Faithful Natural Language Explanation
por: Bandyopadhyay, Dibyanayan, et al.
Publicado: (2025)
por: Bandyopadhyay, Dibyanayan, et al.
Publicado: (2025)
Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach
por: Wojciechowski, Adam, et al.
Publicado: (2024)
por: Wojciechowski, Adam, et al.
Publicado: (2024)
Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language Models
por: Yeo, Wei Jie, et al.
Publicado: (2024)
por: Yeo, Wei Jie, et al.
Publicado: (2024)
Local Explanations and Self-Explanations for Assessing Faithfulness in black-box LLMs
por: Fragkathoulas, Christos, et al.
Publicado: (2024)
por: Fragkathoulas, Christos, et al.
Publicado: (2024)
FaithLM: Towards Faithful Explanations for Large Language Models
por: Chuang, Yu-Neng, et al.
Publicado: (2024)
por: Chuang, Yu-Neng, et al.
Publicado: (2024)
Cross-Refine: Improving Natural Language Explanation Generation by Learning in Tandem
por: Wang, Qianli, et al.
Publicado: (2024)
por: Wang, Qianli, et al.
Publicado: (2024)
Verification and Refinement of Natural Language Explanations through LLM-Symbolic Theorem Proving
por: Quan, Xin, et al.
Publicado: (2024)
por: Quan, Xin, et al.
Publicado: (2024)
DYNAMICQA: Tracing Internal Knowledge Conflicts in Language Models
por: Marjanović, Sara Vera, et al.
Publicado: (2024)
por: Marjanović, Sara Vera, et al.
Publicado: (2024)
Self-Critique-Guided Curiosity Refinement: Enhancing Honesty and Helpfulness in Large Language Models via In-Context Learning
por: Ho, Duc Hieu, et al.
Publicado: (2025)
por: Ho, Duc Hieu, et al.
Publicado: (2025)
MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
por: Liu, Gabrielle Kaili-May, et al.
Publicado: (2025)
por: Liu, Gabrielle Kaili-May, et al.
Publicado: (2025)
Grounding Natural Language to SQL Translation with Data-Based Self-Explanations
por: Fan, Yuankai, et al.
Publicado: (2024)
por: Fan, Yuankai, et al.
Publicado: (2024)
Causally Grounded Mechanistic Interpretability for LLMs with Faithful Natural-Language Explanations
por: Mahale, Ajay Pravin
Publicado: (2026)
por: Mahale, Ajay Pravin
Publicado: (2026)
Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
por: Agarwal, Chirag, et al.
Publicado: (2024)
por: Agarwal, Chirag, et al.
Publicado: (2024)
Illocutionary Explanation Planning for Source-Faithful Explanations in Retrieval-Augmented Language Models
por: Sovrano, Francesco, et al.
Publicado: (2026)
por: Sovrano, Francesco, et al.
Publicado: (2026)
Situated Natural Language Explanations
por: Zhu, Zining, et al.
Publicado: (2023)
por: Zhu, Zining, et al.
Publicado: (2023)
Faithful-Patchscopes: Understanding and Mitigating Model Bias in Hidden Representations Explanation of Large Language Models
por: Gong, Xilin, et al.
Publicado: (2026)
por: Gong, Xilin, et al.
Publicado: (2026)
Towards Faithful Explanations for Text Classification with Robustness Improvement and Explanation Guided Training
por: Li, Dongfang, et al.
Publicado: (2023)
por: Li, Dongfang, et al.
Publicado: (2023)
Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning
por: Zhu, Chenghao, et al.
Publicado: (2025)
por: Zhu, Chenghao, et al.
Publicado: (2025)
Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
por: Siegel, Noah Y., et al.
Publicado: (2025)
por: Siegel, Noah Y., et al.
Publicado: (2025)
The Probabilities Also Matter: A More Faithful Metric for Faithfulness of Free-Text Explanations in Large Language Models
por: Siegel, Noah Y., et al.
Publicado: (2024)
por: Siegel, Noah Y., et al.
Publicado: (2024)
RefineCoder: Iterative Improving of Large Language Models via Adaptive Critique Refinement for Code Generation
por: Zhou, Changzhi, et al.
Publicado: (2025)
por: Zhou, Changzhi, et al.
Publicado: (2025)
Towards Faithful Model Explanation in NLP: A Survey
por: Lyu, Qing, et al.
Publicado: (2022)
por: Lyu, Qing, et al.
Publicado: (2022)
Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations
por: Matton, Katie, et al.
Publicado: (2025)
por: Matton, Katie, et al.
Publicado: (2025)
Digital Socrates: Evaluating LLMs through Explanation Critiques
por: Gu, Yuling, et al.
Publicado: (2023)
por: Gu, Yuling, et al.
Publicado: (2023)
Large Language Model Agents Are Not Always Faithful Self-Evolvers
por: Zhao, Weixiang, et al.
Publicado: (2026)
por: Zhao, Weixiang, et al.
Publicado: (2026)
Faithfulness and the Notion of Adversarial Sensitivity in NLP Explanations
por: Manna, Supriya, et al.
Publicado: (2024)
por: Manna, Supriya, et al.
Publicado: (2024)
Ejemplares similares
-
Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
por: Sun, Jingyi, et al.
Publicado: (2024) -
Evaluation Framework for Highlight Explanations of Context Utilisation in Language Models
por: Sun, Jingyi, et al.
Publicado: (2025) -
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
por: Sun, Jingyi, et al.
Publicado: (2026) -
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
por: Wang, Qianli, et al.
Publicado: (2026) -
Learning from Self Critique and Refinement for Faithful LLM Summarization
por: Hu, Ting-Yao, et al.
Publicado: (2025)