Concept-based explainability for an EEG transformer model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gjølbye, Anders, Lehn-Schiøler, William, Jónsdóttir, Áshildur, Arnardóttir, Bergdís, Hansen, Lars Kai
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910963386548224
author Gjølbye, Anders
Lehn-Schiøler, William
Jónsdóttir, Áshildur
Arnardóttir, Bergdís
Hansen, Lars Kai
author_facet Gjølbye, Anders
Lehn-Schiøler, William
Jónsdóttir, Áshildur
Arnardóttir, Bergdís
Hansen, Lars Kai
contents Deep learning models are complex due to their size, structure, and inherent randomness in training procedures. Additional complexity arises from the selection of datasets and inductive biases. Addressing these challenges for explainability, Kim et al. (2018) introduced Concept Activation Vectors (CAVs), which aim to understand deep models' internal states in terms of human-aligned concepts. These concepts correspond to directions in latent space, identified using linear discriminants. Although this method was first applied to image classification, it was later adapted to other domains, including natural language processing. In this work, we attempt to apply the method to electroencephalogram (EEG) data for explainability in Kostas et al.'s BENDR (2021), a large-scale transformer model. A crucial part of this endeavor involves defining the explanatory concepts and selecting relevant datasets to ground concepts in the latent space. Our focus is on two mechanisms for EEG concept formation: the use of externally labeled EEG datasets, and the application of anatomically defined concepts. The former approach is a straightforward generalization of methods used in image classification, while the latter is novel and specific to EEG. We present evidence that both approaches to concept formation yield valuable insights into the representations learned by deep EEG models.
format Preprint
id arxiv_https___arxiv_org_abs_2307_12745
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Concept-based explainability for an EEG transformer model
Gjølbye, Anders
Lehn-Schiøler, William
Jónsdóttir, Áshildur
Arnardóttir, Bergdís
Hansen, Lars Kai
Machine Learning
Signal Processing
Deep learning models are complex due to their size, structure, and inherent randomness in training procedures. Additional complexity arises from the selection of datasets and inductive biases. Addressing these challenges for explainability, Kim et al. (2018) introduced Concept Activation Vectors (CAVs), which aim to understand deep models' internal states in terms of human-aligned concepts. These concepts correspond to directions in latent space, identified using linear discriminants. Although this method was first applied to image classification, it was later adapted to other domains, including natural language processing. In this work, we attempt to apply the method to electroencephalogram (EEG) data for explainability in Kostas et al.'s BENDR (2021), a large-scale transformer model. A crucial part of this endeavor involves defining the explanatory concepts and selecting relevant datasets to ground concepts in the latent space. Our focus is on two mechanisms for EEG concept formation: the use of externally labeled EEG datasets, and the application of anatomically defined concepts. The former approach is a straightforward generalization of methods used in image classification, while the latter is novel and specific to EEG. We present evidence that both approaches to concept formation yield valuable insights into the representations learned by deep EEG models.
title Concept-based explainability for an EEG transformer model
topic Machine Learning
Signal Processing
url https://arxiv.org/abs/2307.12745