Faithful and Robust Local Interpretability for Textual Predictions

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lopardo, Gianluigi, Precioso, Frederic, Garreau, Damien
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914745554042880
author Lopardo, Gianluigi
Precioso, Frederic
Garreau, Damien
author_facet Lopardo, Gianluigi
Precioso, Frederic
Garreau, Damien
contents Interpretability is essential for machine learning models to be trusted and deployed in critical domains. However, existing methods for interpreting text models are often complex, lack mathematical foundations, and their performance is not guaranteed. In this paper, we propose FRED (Faithful and Robust Explainer for textual Documents), a novel method for interpreting predictions over text. FRED offers three key insights to explain a model prediction: (1) it identifies the minimal set of words in a document whose removal has the strongest influence on the prediction, (2) it assigns an importance score to each token, reflecting its influence on the model's output, and (3) it provides counterfactual explanations by generating examples similar to the original document, but leading to a different prediction. We establish the reliability of FRED through formal definitions and theoretical analyses on interpretable classifiers. Additionally, our empirical evaluation against state-of-the-art methods demonstrates the effectiveness of FRED in providing insights into text models.
format Preprint
id arxiv_https___arxiv_org_abs_2311_01605
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Faithful and Robust Local Interpretability for Textual Predictions
Lopardo, Gianluigi
Precioso, Frederic
Garreau, Damien
Computation and Language
Machine Learning
Interpretability is essential for machine learning models to be trusted and deployed in critical domains. However, existing methods for interpreting text models are often complex, lack mathematical foundations, and their performance is not guaranteed. In this paper, we propose FRED (Faithful and Robust Explainer for textual Documents), a novel method for interpreting predictions over text. FRED offers three key insights to explain a model prediction: (1) it identifies the minimal set of words in a document whose removal has the strongest influence on the prediction, (2) it assigns an importance score to each token, reflecting its influence on the model's output, and (3) it provides counterfactual explanations by generating examples similar to the original document, but leading to a different prediction. We establish the reliability of FRED through formal definitions and theoretical analyses on interpretable classifiers. Additionally, our empirical evaluation against state-of-the-art methods demonstrates the effectiveness of FRED in providing insights into text models.
title Faithful and Robust Local Interpretability for Textual Predictions
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2311.01605