Faithful and Robust Local Interpretability for Textual Predictions
Fuente:
arXiv
Saved in:
| Main Authors: | Lopardo, Gianluigi, Precioso, Frederic, Garreau, Damien |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attention Meets Post-hoc Interpretability: A Mathematical Perspective
by: Lopardo, Gianluigi, et al.
Published: (2024)
by: Lopardo, Gianluigi, et al.
Published: (2024)
Understanding Post-hoc Explainers: The Case of Anchors
by: Lopardo, Gianluigi, et al.
Published: (2023)
by: Lopardo, Gianluigi, et al.
Published: (2023)
A Sea of Words: An In-Depth Analysis of Anchors for Text Data
by: Lopardo, Gianluigi, et al.
Published: (2022)
by: Lopardo, Gianluigi, et al.
Published: (2022)
Comparing Feature Importance and Rule Extraction for Interpretability on Text Data
by: Lopardo, Gianluigi, et al.
Published: (2022)
by: Lopardo, Gianluigi, et al.
Published: (2022)
MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations
by: Mitsuzawa, Kensuke, et al.
Published: (2025)
by: Mitsuzawa, Kensuke, et al.
Published: (2025)
Towards Understanding Steering Strength
by: Taimeskhanov, Magamed, et al.
Published: (2026)
by: Taimeskhanov, Magamed, et al.
Published: (2026)
New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing
by: Madsen, Andreas
Published: (2024)
by: Madsen, Andreas
Published: (2024)
When Are Two Scores Better Than One? Investigating Ensembles of Diffusion Models
by: Razafindralambo, Raphaël, et al.
Published: (2026)
by: Razafindralambo, Raphaël, et al.
Published: (2026)
Beyond Mixtures and Products for Ensemble Aggregation: A Likelihood Perspective on Generalized Means
by: Razafindralambo, Raphaël, et al.
Published: (2026)
by: Razafindralambo, Raphaël, et al.
Published: (2026)
Geometry-Adaptive Explainer for Faithful Dictionary-Based Interpretability under Distribution Shift
by: Lim, Sungjun, et al.
Published: (2026)
by: Lim, Sungjun, et al.
Published: (2026)
WolBanking77: Wolof Banking Speech Intent Classification Dataset
by: Kandji, Abdou Karim, et al.
Published: (2025)
by: Kandji, Abdou Karim, et al.
Published: (2025)
Transformer Circuit Faithfulness Metrics are not Robust
by: Miller, Joseph, et al.
Published: (2024)
by: Miller, Joseph, et al.
Published: (2024)
Robust Infidelity: When Faithfulness Measures on Masked Language Models Are Misleading
by: Crothers, Evan, et al.
Published: (2023)
by: Crothers, Evan, et al.
Published: (2023)
Towards Faithful and Robust LLM Specialists for Evidence-Based Question-Answering
by: Schimanski, Tobias, et al.
Published: (2024)
by: Schimanski, Tobias, et al.
Published: (2024)
OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction
by: Hemadri, Raghu Vamshi, et al.
Published: (2025)
by: Hemadri, Raghu Vamshi, et al.
Published: (2025)
MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
by: Liu, Gabrielle Kaili-May, et al.
Published: (2025)
by: Liu, Gabrielle Kaili-May, et al.
Published: (2025)
Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning
by: Jia, Jinghan, et al.
Published: (2026)
by: Jia, Jinghan, et al.
Published: (2026)
Faithfulness Measurable Masked Language Models
by: Madsen, Andreas, et al.
Published: (2023)
by: Madsen, Andreas, et al.
Published: (2023)
Mapping Faithful Reasoning in Language Models
by: Li, Jiazheng, et al.
Published: (2025)
by: Li, Jiazheng, et al.
Published: (2025)
Are LLM Decisions Faithful to Verbal Confidence?
by: Wang, Jiawei, et al.
Published: (2026)
by: Wang, Jiawei, et al.
Published: (2026)
Multilingual Self-Taught Faithfulness Evaluators
by: Alfano, Carlo, et al.
Published: (2025)
by: Alfano, Carlo, et al.
Published: (2025)
Understanding Textual Emotion Through Emoji Prediction
by: Gordon, Ethan, et al.
Published: (2025)
by: Gordon, Ethan, et al.
Published: (2025)
FaithLM: Towards Faithful Explanations for Large Language Models
by: Chuang, Yu-Neng, et al.
Published: (2024)
by: Chuang, Yu-Neng, et al.
Published: (2024)
The Risks of Recourse in Binary Classification
by: Fokkema, Hidde, et al.
Published: (2023)
by: Fokkema, Hidde, et al.
Published: (2023)
MPAT: Building Robust Deep Neural Networks against Textual Adversarial Attacks
by: Zhang, Fangyuan, et al.
Published: (2024)
by: Zhang, Fangyuan, et al.
Published: (2024)
CAM-Based Methods Can See through Walls
by: Taimeskhanov, Magamed, et al.
Published: (2024)
by: Taimeskhanov, Magamed, et al.
Published: (2024)
GLEAMS: Bridging the Gap Between Local and Global Explanations
by: Visani, Giorgio, et al.
Published: (2024)
by: Visani, Giorgio, et al.
Published: (2024)
Feature Attribution from First Principles
by: Taimeskhanov, Magamed, et al.
Published: (2025)
by: Taimeskhanov, Magamed, et al.
Published: (2025)
On The Variability of Concept Activation Vectors
by: Wenkmann, Julia, et al.
Published: (2025)
by: Wenkmann, Julia, et al.
Published: (2025)
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
by: Zhang, Xinyu, et al.
Published: (2023)
by: Zhang, Xinyu, et al.
Published: (2023)
LLMForecaster: Improving Seasonal Event Forecasts with Unstructured Textual Data
by: Zhang, Hanyu, et al.
Published: (2024)
by: Zhang, Hanyu, et al.
Published: (2024)
Textual Gradients are a Flawed Metaphor for Automatic Prompt Optimization
by: Melcer, Daniel, et al.
Published: (2025)
by: Melcer, Daniel, et al.
Published: (2025)
GraphNarrator: Generating Textual Explanations for Graph Neural Networks
by: Pan, Bo, et al.
Published: (2024)
by: Pan, Bo, et al.
Published: (2024)
GenFighter: A Generative and Evolutive Textual Attack Removal
by: Islam, Md Athikul, et al.
Published: (2024)
by: Islam, Md Athikul, et al.
Published: (2024)
RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
by: Yu, Zichun, et al.
Published: (2025)
by: Yu, Zichun, et al.
Published: (2025)
Retrieval-Augmented and Knowledge-Grounded Language Models for Faithful Clinical Medicine
by: Liu, Fenglin, et al.
Published: (2022)
by: Liu, Fenglin, et al.
Published: (2022)
FaithfulSAE: Towards Capturing Faithful Features with Sparse Autoencoders without External Dataset Dependencies
by: Cho, Seonglae, et al.
Published: (2025)
by: Cho, Seonglae, et al.
Published: (2025)
A Curious Case of Searching for the Correlation between Training Data and Adversarial Robustness of Transformer Textual Models
by: Dang, Cuong, et al.
Published: (2024)
by: Dang, Cuong, et al.
Published: (2024)
TIDE: Textual Identity Detection for Evaluating and Augmenting Classification and Language Models
by: Klu, Emmanuel, et al.
Published: (2023)
by: Klu, Emmanuel, et al.
Published: (2023)
TRPrompt: Bootstrapping Query-Aware Prompt Optimization from Textual Rewards
by: Nica, Andreea, et al.
Published: (2025)
by: Nica, Andreea, et al.
Published: (2025)
Similar Items
-
Attention Meets Post-hoc Interpretability: A Mathematical Perspective
by: Lopardo, Gianluigi, et al.
Published: (2024) -
Understanding Post-hoc Explainers: The Case of Anchors
by: Lopardo, Gianluigi, et al.
Published: (2023) -
A Sea of Words: An In-Depth Analysis of Anchors for Text Data
by: Lopardo, Gianluigi, et al.
Published: (2022) -
Comparing Feature Importance and Rule Extraction for Interpretability on Text Data
by: Lopardo, Gianluigi, et al.
Published: (2022) -
MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations
by: Mitsuzawa, Kensuke, et al.
Published: (2025)