Explanation Regularisation through the Lens of Attributions
Fuente:
arXiv
Saved in:
| Main Authors: | Ferreira, Pedro, Titov, Ivan, Aziz, Wilker |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
by: Ferreira, Pedro, et al.
Published: (2025)
by: Ferreira, Pedro, et al.
Published: (2025)
Variability Need Not Imply Error: The Case of Adequate but Semantically Distinct Responses
by: Ilia, Evgenia, et al.
Published: (2024)
by: Ilia, Evgenia, et al.
Published: (2024)
Predict the Next Word: Humans exhibit uncertainty in this task and language models _____
by: Ilia, Evgenia, et al.
Published: (2024)
by: Ilia, Evgenia, et al.
Published: (2024)
MedPath: Multi-Domain Cross-Vocabulary Hierarchical Paths for Biomedical Entity Linking
by: Mishra, Nishant, et al.
Published: (2025)
by: Mishra, Nishant, et al.
Published: (2025)
Interpreting Predictive Probabilities: Model Confidence or Human Label Variation?
by: Baan, Joris, et al.
Published: (2024)
by: Baan, Joris, et al.
Published: (2024)
Mitigating Copy Bias in In-Context Learning through Neuron Pruning
by: Ali, Ameen, et al.
Published: (2024)
by: Ali, Ameen, et al.
Published: (2024)
Generalisation First, Memorisation Second? Memorisation Localisation for Natural Language Classification Tasks
by: Dankers, Verna, et al.
Published: (2024)
by: Dankers, Verna, et al.
Published: (2024)
Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
by: Baan, Joris, et al.
Published: (2026)
by: Baan, Joris, et al.
Published: (2026)
Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations
by: Hong, Pingjun, et al.
Published: (2025)
by: Hong, Pingjun, et al.
Published: (2025)
Unlearning Traces the Influential Training Data of Language Models
by: Isonuma, Masaru, et al.
Published: (2024)
by: Isonuma, Masaru, et al.
Published: (2024)
M-Wanda: Improving One-Shot Pruning for Multilingual LLMs
by: Choenni, Rochelle, et al.
Published: (2025)
by: Choenni, Rochelle, et al.
Published: (2025)
Learning from Sufficient Rationales: Analysing the Relationship Between Explanation Faithfulness and Token-level Regularisation Strategies
by: Kamp, Jonathan, et al.
Published: (2025)
by: Kamp, Jonathan, et al.
Published: (2025)
Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection
by: Ramírez, Guillem, et al.
Published: (2024)
by: Ramírez, Guillem, et al.
Published: (2024)
Strengthening Structural Inductive Biases by Pre-training to Perform Syntactic Transformations
by: Lindemann, Matthias, et al.
Published: (2024)
by: Lindemann, Matthias, et al.
Published: (2024)
SIP: Injecting a Structural Inductive Bias into a Seq2Seq Model by Simulation
by: Lindemann, Matthias, et al.
Published: (2023)
by: Lindemann, Matthias, et al.
Published: (2023)
ChartLens: Fine-grained Visual Attribution in Charts
by: Suri, Manan, et al.
Published: (2025)
by: Suri, Manan, et al.
Published: (2025)
What's New in My Data? Novelty Exploration via Contrastive Generation
by: Isonuma, Masaru, et al.
Published: (2024)
by: Isonuma, Masaru, et al.
Published: (2024)
Teaching Language Models to Faithfully Express their Uncertainty
by: Eikema, Bryan, et al.
Published: (2025)
by: Eikema, Bryan, et al.
Published: (2025)
Controlling What You Share: Assessing Language Model Adherence to Privacy Preferences
by: Ramírez, Guillem, et al.
Published: (2025)
by: Ramírez, Guillem, et al.
Published: (2025)
Belief Attribution as Mental Explanation: The Role of Accuracy, Informativity, and Causality
by: Ying, Lance, et al.
Published: (2025)
by: Ying, Lance, et al.
Published: (2025)
Evaluating Evidence Attribution in Generated Fact Checking Explanations
by: Xing, Rui, et al.
Published: (2024)
by: Xing, Rui, et al.
Published: (2024)
Shared Doubt: Zero-shot Cross-Lingual Confidence Estimation for Language Models
by: Kyriakou, Athina, et al.
Published: (2026)
by: Kyriakou, Athina, et al.
Published: (2026)
Joint Localization and Activation Editing for Low-Resource Fine-Tuning
by: Lai, Wen, et al.
Published: (2025)
by: Lai, Wen, et al.
Published: (2025)
Tox-BART: Leveraging Toxicity Attributes for Explanation Generation of Implicit Hate Speech
by: Yadav, Neemesh, et al.
Published: (2024)
by: Yadav, Neemesh, et al.
Published: (2024)
Cache & Distil: Optimising API Calls to Large Language Models
by: Ramírez, Guillem, et al.
Published: (2023)
by: Ramírez, Guillem, et al.
Published: (2023)
Anthropomimetic Uncertainty: What Verbalized Uncertainty in Language Models is Missing
by: Ulmer, Dennis, et al.
Published: (2025)
by: Ulmer, Dennis, et al.
Published: (2025)
Detecting and Pruning Prominent but Detrimental Neurons in Large Language Models
by: Ali, Ameen, et al.
Published: (2025)
by: Ali, Ameen, et al.
Published: (2025)
Not All Explanations Simulate Equally: Comparing Verbalized Feature Attributions and Self-Generated Rationales
by: Hong, Pingjun, et al.
Published: (2026)
by: Hong, Pingjun, et al.
Published: (2026)
CARMA: Enhanced Compositionality in LLMs via Advanced Regularisation and Mutual Information Alignment
by: Aljaafari, Nura, et al.
Published: (2025)
by: Aljaafari, Nura, et al.
Published: (2025)
Memorization in Language Models through the Lens of Intrinsic Dimension
by: Arnold, Stefan
Published: (2025)
by: Arnold, Stefan
Published: (2025)
Enhancing Long Document Long Form Summarisation with Self-Planning
by: Du, Xiaotang, et al.
Published: (2025)
by: Du, Xiaotang, et al.
Published: (2025)
Analysing Cross-Speaker Convergence in Face-to-Face Dialogue through the Lens of Automatically Detected Shared Linguistic Constructions
by: Ghaleb, Esam, et al.
Published: (2024)
by: Ghaleb, Esam, et al.
Published: (2024)
Analyzing Wrap-Up Effects through an Information-Theoretic Lens
by: Meister, Clara, et al.
Published: (2022)
by: Meister, Clara, et al.
Published: (2022)
LLM Bias Detection and Mitigation through the Lens of Desired Distributions
by: Shrestha, Ingroj, et al.
Published: (2025)
by: Shrestha, Ingroj, et al.
Published: (2025)
Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights
by: Karthika, N J, et al.
Published: (2025)
by: Karthika, N J, et al.
Published: (2025)
On the Interplay between Musical Preferences and Personality through the Lens of Language
by: Shem-Tov, Eliran, et al.
Published: (2025)
by: Shem-Tov, Eliran, et al.
Published: (2025)
Finding Culture-Sensitive Neurons in Vision-Language Models
by: Zhao, Xiutian, et al.
Published: (2025)
by: Zhao, Xiutian, et al.
Published: (2025)
Is Modularity Transferable? A Case Study through the Lens of Knowledge Distillation
by: Klimaszewski, Mateusz, et al.
Published: (2024)
by: Klimaszewski, Mateusz, et al.
Published: (2024)
Analyzing Feed-Forward Blocks in Transformers through the Lens of Attention Maps
by: Kobayashi, Goro, et al.
Published: (2023)
by: Kobayashi, Goro, et al.
Published: (2023)
Analyzing LLM Instruction Optimization for Tabular Fact Verification
by: Du, Xiaotang, et al.
Published: (2026)
by: Du, Xiaotang, et al.
Published: (2026)
Similar Items
-
Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
by: Ferreira, Pedro, et al.
Published: (2025) -
Variability Need Not Imply Error: The Case of Adequate but Semantically Distinct Responses
by: Ilia, Evgenia, et al.
Published: (2024) -
Predict the Next Word: Humans exhibit uncertainty in this task and language models _____
by: Ilia, Evgenia, et al.
Published: (2024) -
MedPath: Multi-Domain Cross-Vocabulary Hierarchical Paths for Biomedical Entity Linking
by: Mishra, Nishant, et al.
Published: (2025) -
Interpreting Predictive Probabilities: Model Confidence or Human Label Variation?
by: Baan, Joris, et al.
Published: (2024)