Linear Explanations for Individual Neurons

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Oikarinen, Tuomas, Weng, Tsui-Wei
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929339623276544
author Oikarinen, Tuomas
Weng, Tsui-Wei
author_facet Oikarinen, Tuomas
Weng, Tsui-Wei
contents In recent years many methods have been developed to understand the internal workings of neural networks, often by describing the function of individual neurons in the model. However, these methods typically only focus on explaining the very highest activations of a neuron. In this paper we show this is not sufficient, and that the highest activation range is only responsible for a very small percentage of the neuron's causal effect. In addition, inputs causing lower activations are often very different and can't be reliably predicted by only looking at high activations. We propose that neurons should instead be understood as a linear combination of concepts, and develop an efficient method for producing these linear explanations. In addition, we show how to automatically evaluate description quality using simulation, i.e. predicting neuron activations on unseen inputs in vision setting.
format Preprint
id arxiv_https___arxiv_org_abs_2405_06855
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Linear Explanations for Individual Neurons
Oikarinen, Tuomas
Weng, Tsui-Wei
Machine Learning
Computer Vision and Pattern Recognition
In recent years many methods have been developed to understand the internal workings of neural networks, often by describing the function of individual neurons in the model. However, these methods typically only focus on explaining the very highest activations of a neuron. In this paper we show this is not sufficient, and that the highest activation range is only responsible for a very small percentage of the neuron's causal effect. In addition, inputs causing lower activations are often very different and can't be reliably predicted by only looking at high activations. We propose that neurons should instead be understood as a linear combination of concepts, and develop an efficient method for producing these linear explanations. In addition, we show how to automatically evaluate description quality using simulation, i.e. predicting neuron activations on unseen inputs in vision setting.
title Linear Explanations for Individual Neurons
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.06855