CoSy: Evaluating Textual Explanations of Neurons

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kopf, Laura, Bommer, Philine Lou, Hedström, Anna, Lapuschkin, Sebastian, Höhne, Marina M. -C., Bykov, Kirill
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912144444882944
author Kopf, Laura
Bommer, Philine Lou
Hedström, Anna
Lapuschkin, Sebastian
Höhne, Marina M. -C.
Bykov, Kirill
author_facet Kopf, Laura
Bommer, Philine Lou
Hedström, Anna
Lapuschkin, Sebastian
Höhne, Marina M. -C.
Bykov, Kirill
contents A crucial aspect of understanding the complex nature of Deep Neural Networks (DNNs) is the ability to explain learned concepts within their latent representations. While methods exist to connect neurons to human-understandable textual descriptions, evaluating the quality of these explanations is challenging due to the lack of a unified quantitative approach. We introduce CoSy (Concept Synthesis), a novel, architecture-agnostic framework for evaluating textual explanations of latent neurons. Given textual explanations, our proposed framework uses a generative model conditioned on textual input to create data points representing the explanations. By comparing the neuron's response to these generated data points and control data points, we can estimate the quality of the explanation. We validate our framework through sanity checks and benchmark various neuron description methods for Computer Vision tasks, revealing significant differences in quality.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20331
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CoSy: Evaluating Textual Explanations of Neurons
Kopf, Laura
Bommer, Philine Lou
Hedström, Anna
Lapuschkin, Sebastian
Höhne, Marina M. -C.
Bykov, Kirill
Machine Learning
Artificial Intelligence
Computation and Language
A crucial aspect of understanding the complex nature of Deep Neural Networks (DNNs) is the ability to explain learned concepts within their latent representations. While methods exist to connect neurons to human-understandable textual descriptions, evaluating the quality of these explanations is challenging due to the lack of a unified quantitative approach. We introduce CoSy (Concept Synthesis), a novel, architecture-agnostic framework for evaluating textual explanations of latent neurons. Given textual explanations, our proposed framework uses a generative model conditioned on textual input to create data points representing the explanations. By comparing the neuron's response to these generated data points and control data points, we can estimate the quality of the explanation. We validate our framework through sanity checks and benchmark various neuron description methods for Computer Vision tasks, revealing significant differences in quality.
title CoSy: Evaluating Textual Explanations of Neurons
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2405.20331