Obtaining Example-Based Explanations from Deep Neural Networks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Dong, Genghua, Boström, Henrik, Vazirgiannis, Michalis, Bresson, Roman
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912250048020480
author Dong, Genghua
Boström, Henrik
Vazirgiannis, Michalis
Bresson, Roman
author_facet Dong, Genghua
Boström, Henrik
Vazirgiannis, Michalis
Bresson, Roman
contents Most techniques for explainable machine learning focus on feature attribution, i.e., values are assigned to the features such that their sum equals the prediction. Example attribution is another form of explanation that assigns weights to the training examples, such that their scalar product with the labels equals the prediction. The latter may provide valuable complementary information to feature attribution, in particular in cases where the features are not easily interpretable. Current example-based explanation techniques have targeted a few model types only, such as k-nearest neighbors and random forests. In this work, a technique for obtaining example-based explanations from deep neural networks (EBE-DNN) is proposed. The basic idea is to use the deep neural network to obtain an embedding, which is employed by a k-nearest neighbor classifier to form a prediction; the example attribution can hence straightforwardly be derived from the latter. Results from an empirical investigation show that EBE-DNN can provide highly concentrated example attributions, i.e., the predictions can be explained with few training examples, without reducing accuracy compared to the original deep neural network. Another important finding from the empirical investigation is that the choice of layer to use for the embeddings may have a large impact on the resulting accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19768
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Obtaining Example-Based Explanations from Deep Neural Networks
Dong, Genghua
Boström, Henrik
Vazirgiannis, Michalis
Bresson, Roman
Machine Learning
Artificial Intelligence
Most techniques for explainable machine learning focus on feature attribution, i.e., values are assigned to the features such that their sum equals the prediction. Example attribution is another form of explanation that assigns weights to the training examples, such that their scalar product with the labels equals the prediction. The latter may provide valuable complementary information to feature attribution, in particular in cases where the features are not easily interpretable. Current example-based explanation techniques have targeted a few model types only, such as k-nearest neighbors and random forests. In this work, a technique for obtaining example-based explanations from deep neural networks (EBE-DNN) is proposed. The basic idea is to use the deep neural network to obtain an embedding, which is employed by a k-nearest neighbor classifier to form a prediction; the example attribution can hence straightforwardly be derived from the latter. Results from an empirical investigation show that EBE-DNN can provide highly concentrated example attributions, i.e., the predictions can be explained with few training examples, without reducing accuracy compared to the original deep neural network. Another important finding from the empirical investigation is that the choice of layer to use for the embeddings may have a large impact on the resulting accuracy.
title Obtaining Example-Based Explanations from Deep Neural Networks
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2502.19768