Neuron-based explanations of neural networks sacrifice completeness and interpretability

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Dey, Nolan, Taylor, Eric, Wong, Alexander, Tripp, Bryan, Taylor, Graham W.
Formato: Preprint
Publicado: 2020
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916657010573312
author Dey, Nolan
Taylor, Eric
Wong, Alexander
Tripp, Bryan
Taylor, Graham W.
author_facet Dey, Nolan
Taylor, Eric
Wong, Alexander
Tripp, Bryan
Taylor, Graham W.
contents High quality explanations of neural networks (NNs) should exhibit two key properties. Completeness ensures that they accurately reflect a network's function and interpretability makes them understandable to humans. Many existing methods provide explanations of individual neurons within a network. In this work we provide evidence that for AlexNet pretrained on ImageNet, neuron-based explanation methods sacrifice both completeness and interpretability compared to activation principal components. Neurons are a poor basis for AlexNet embeddings because they don't account for the distributed nature of these representations. By examining two quantitative measures of completeness and conducting a user study to measure interpretability, we show the most important principal components provide more complete and interpretable explanations than the most important neurons. Much of the activation variance may be explained by examining relatively few high-variance PCs, as opposed to studying every neuron. These principal components also strongly affect network function, and are significantly more interpretable than neurons. Our findings suggest that explanation methods for networks like AlexNet should avoid using neurons as a basis for embeddings and instead choose a basis, such as principal components, which accounts for the high dimensional and distributed nature of a network's internal representations. Interactive demo and code available at https://ndey96.github.io/neuron-explanations-sacrifice.
format Preprint
id arxiv_https___arxiv_org_abs_2011_03043
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle Neuron-based explanations of neural networks sacrifice completeness and interpretability
Dey, Nolan
Taylor, Eric
Wong, Alexander
Tripp, Bryan
Taylor, Graham W.
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.10
High quality explanations of neural networks (NNs) should exhibit two key properties. Completeness ensures that they accurately reflect a network's function and interpretability makes them understandable to humans. Many existing methods provide explanations of individual neurons within a network. In this work we provide evidence that for AlexNet pretrained on ImageNet, neuron-based explanation methods sacrifice both completeness and interpretability compared to activation principal components. Neurons are a poor basis for AlexNet embeddings because they don't account for the distributed nature of these representations. By examining two quantitative measures of completeness and conducting a user study to measure interpretability, we show the most important principal components provide more complete and interpretable explanations than the most important neurons. Much of the activation variance may be explained by examining relatively few high-variance PCs, as opposed to studying every neuron. These principal components also strongly affect network function, and are significantly more interpretable than neurons. Our findings suggest that explanation methods for networks like AlexNet should avoid using neurons as a basis for embeddings and instead choose a basis, such as principal components, which accounts for the high dimensional and distributed nature of a network's internal representations. Interactive demo and code available at https://ndey96.github.io/neuron-explanations-sacrifice.
title Neuron-based explanations of neural networks sacrifice completeness and interpretability
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
I.2.10
url https://arxiv.org/abs/2011.03043