Not All Explanations Simulate Equally: Comparing Verbalized Feature Attributions and Self-Generated Rationales
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hong, Pingjun, Roth, Benjamin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Do LLM Self-Explanations Help Users Predict Model Behavior? Evaluating Counterfactual Simulatability with Pragmatic Perturbations
von: Hong, Pingjun, et al.
Veröffentlicht: (2026)
von: Hong, Pingjun, et al.
Veröffentlicht: (2026)
Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization
von: Chen, Beiduo, et al.
Veröffentlicht: (2026)
von: Chen, Beiduo, et al.
Veröffentlicht: (2026)
Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
von: Hong, Pingjun, et al.
Veröffentlicht: (2025)
von: Hong, Pingjun, et al.
Veröffentlicht: (2025)
Not All Data Are Unlearned Equally
von: Krishnan, Aravind, et al.
Veröffentlicht: (2025)
von: Krishnan, Aravind, et al.
Veröffentlicht: (2025)
Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs
von: Xia, Yuxi, et al.
Veröffentlicht: (2026)
von: Xia, Yuxi, et al.
Veröffentlicht: (2026)
LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inference
von: Hong, Pingjun, et al.
Veröffentlicht: (2025)
von: Hong, Pingjun, et al.
Veröffentlicht: (2025)
Not All Demonstration Examples are Equally Beneficial: Reweighting Demonstration Examples for In-Context Learning
von: Yang, Zhe, et al.
Veröffentlicht: (2023)
von: Yang, Zhe, et al.
Veröffentlicht: (2023)
Keyword-Centric Prompting for One-Shot Event Detection with Self-Generated Rationale Enhancements
von: Li, Ziheng, et al.
Veröffentlicht: (2025)
von: Li, Ziheng, et al.
Veröffentlicht: (2025)
An Evaluation of Explanation Methods for Black-Box Detectors of Machine-Generated Text
von: Schoenegger, Loris, et al.
Veröffentlicht: (2024)
von: Schoenegger, Loris, et al.
Veröffentlicht: (2024)
Compact Example-Based Explanations for Language Models
von: Schoenegger, Loris, et al.
Veröffentlicht: (2026)
von: Schoenegger, Loris, et al.
Veröffentlicht: (2026)
Fine-Grained Perspectives: Modeling Explanations with Annotator-Specific Rationales
von: Sarumi, Olufunke O., et al.
Veröffentlicht: (2026)
von: Sarumi, Olufunke O., et al.
Veröffentlicht: (2026)
Do Large Language Models Speak All Languages Equally? A Comparative Study in Low-Resource Settings
von: Hasan, Md. Arid, et al.
Veröffentlicht: (2024)
von: Hasan, Md. Arid, et al.
Veröffentlicht: (2024)
Calibrating Verbalized Confidence with Self-Generated Distractors
von: Wang, Victor, et al.
Veröffentlicht: (2025)
von: Wang, Victor, et al.
Veröffentlicht: (2025)
Self-Explanation in Social AI Agents
von: Basappa, Rhea, et al.
Veröffentlicht: (2025)
von: Basappa, Rhea, et al.
Veröffentlicht: (2025)
Evaluating Evidence Attribution in Generated Fact Checking Explanations
von: Xing, Rui, et al.
Veröffentlicht: (2024)
von: Xing, Rui, et al.
Veröffentlicht: (2024)
Explanation Regularisation through the Lens of Attributions
von: Ferreira, Pedro, et al.
Veröffentlicht: (2024)
von: Ferreira, Pedro, et al.
Veröffentlicht: (2024)
Neologism Learning for Controllability and Self-Verbalization
von: Hewitt, John, et al.
Veröffentlicht: (2025)
von: Hewitt, John, et al.
Veröffentlicht: (2025)
Verbalized Algorithms: Classical Algorithms are All You Need (Mostly)
von: Lall, Supriya, et al.
Veröffentlicht: (2025)
von: Lall, Supriya, et al.
Veröffentlicht: (2025)
Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations
von: Hinterleitner, Lukas, et al.
Veröffentlicht: (2026)
von: Hinterleitner, Lukas, et al.
Veröffentlicht: (2026)
Not All Tokens Matter Equally: Dynamic In-context Vector Distillation with Decisive-Token Supervision for Long-form Medical Report Generation
von: Wu, Ning, et al.
Veröffentlicht: (2026)
von: Wu, Ning, et al.
Veröffentlicht: (2026)
Self-Explaining Hate Speech Detection with Moral Rationales
von: Vargas, Francielle, et al.
Veröffentlicht: (2026)
von: Vargas, Francielle, et al.
Veröffentlicht: (2026)
Rationale-Aware Answer Verification by Pairwise Self-Evaluation
von: Kawabata, Akira, et al.
Veröffentlicht: (2024)
von: Kawabata, Akira, et al.
Veröffentlicht: (2024)
Tox-BART: Leveraging Toxicity Attributes for Explanation Generation of Implicit Hate Speech
von: Yadav, Neemesh, et al.
Veröffentlicht: (2024)
von: Yadav, Neemesh, et al.
Veröffentlicht: (2024)
Evaluating Large Language Models for Cross-Lingual Retrieval
von: Zuo, Longfei, et al.
Veröffentlicht: (2025)
von: Zuo, Longfei, et al.
Veröffentlicht: (2025)
Towards Efficient CoT Distillation: Self-Guided Rationale Selector for Better Performance with Fewer Rationales
von: Yan, Jianzhi, et al.
Veröffentlicht: (2025)
von: Yan, Jianzhi, et al.
Veröffentlicht: (2025)
CLSGen: A Dual-Head Fine-Tuning Framework for Joint Probabilistic Classification and Verbalized Explanation
von: Yoon, WonJin, et al.
Veröffentlicht: (2026)
von: Yoon, WonJin, et al.
Veröffentlicht: (2026)
Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations
von: Ji, Ziwei, et al.
Veröffentlicht: (2025)
von: Ji, Ziwei, et al.
Veröffentlicht: (2025)
Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation
von: Zhao, Haiyan, et al.
Veröffentlicht: (2026)
von: Zhao, Haiyan, et al.
Veröffentlicht: (2026)
InstructRAG: Instructing Retrieval-Augmented Generation via Self-Synthesized Rationales
von: Wei, Zhepei, et al.
Veröffentlicht: (2024)
von: Wei, Zhepei, et al.
Veröffentlicht: (2024)
RORA: Robust Free-Text Rationale Evaluation
von: Jiang, Zhengping, et al.
Veröffentlicht: (2024)
von: Jiang, Zhengping, et al.
Veröffentlicht: (2024)
Explanation Bias is a Product: Revealing the Hidden Lexical and Position Preferences in Post-Hoc Feature Attribution
von: Kamp, Jonathan, et al.
Veröffentlicht: (2025)
von: Kamp, Jonathan, et al.
Veröffentlicht: (2025)
Explaining Generalization of AI-Generated Text Detectors Through Linguistic Analysis
von: Xia, Yuxi, et al.
Veröffentlicht: (2026)
von: Xia, Yuxi, et al.
Veröffentlicht: (2026)
Belief Attribution as Mental Explanation: The Role of Accuracy, Informativity, and Causality
von: Ying, Lance, et al.
Veröffentlicht: (2025)
von: Ying, Lance, et al.
Veröffentlicht: (2025)
Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability
von: Zhu, Chiwei, et al.
Veröffentlicht: (2025)
von: Zhu, Chiwei, et al.
Veröffentlicht: (2025)
Self-Routing RAG: Binding Selective Retrieval with Knowledge Verbalization
von: Wu, Di, et al.
Veröffentlicht: (2025)
von: Wu, Di, et al.
Veröffentlicht: (2025)
Reinforcing Human Behavior Simulation via Verbal Feedback
von: Sun, Weiwei, et al.
Veröffentlicht: (2026)
von: Sun, Weiwei, et al.
Veröffentlicht: (2026)
Learning from Sufficient Rationales: Analysing the Relationship Between Explanation Faithfulness and Token-level Regularisation Strategies
von: Kamp, Jonathan, et al.
Veröffentlicht: (2025)
von: Kamp, Jonathan, et al.
Veröffentlicht: (2025)
Rationale Behind Essay Scores: Enhancing S-LLM's Multi-Trait Essay Scoring with Rationale Generated by LLMs
von: Chu, SeongYeub, et al.
Veröffentlicht: (2024)
von: Chu, SeongYeub, et al.
Veröffentlicht: (2024)
Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
von: Ferreira, Pedro, et al.
Veröffentlicht: (2025)
von: Ferreira, Pedro, et al.
Veröffentlicht: (2025)
Rationale-Guided Retrieval Augmented Generation for Medical Question Answering
von: Sohn, Jiwoong, et al.
Veröffentlicht: (2024)
von: Sohn, Jiwoong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Do LLM Self-Explanations Help Users Predict Model Behavior? Evaluating Counterfactual Simulatability with Pragmatic Perturbations
von: Hong, Pingjun, et al.
Veröffentlicht: (2026) -
Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization
von: Chen, Beiduo, et al.
Veröffentlicht: (2026) -
Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
von: Hong, Pingjun, et al.
Veröffentlicht: (2025) -
Not All Data Are Unlearned Equally
von: Krishnan, Aravind, et al.
Veröffentlicht: (2025) -
Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs
von: Xia, Yuxi, et al.
Veröffentlicht: (2026)