Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lehn-Schiøler, William, Kjær, Magnus Ruud, Thapa, Rahul, Pedersen, Magnus Guldberg, Storgaard, Anton Mosquera, Williams, Nick, Gatej, Radu, Lehn-Schiøler, Tue, Brink-Kjær, Andreas, Puthusserypady, Sadasivan, Beniczky, Sándor, Zou, James, Hansen, Lars Kai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917524463943680
author Lehn-Schiøler, William
Kjær, Magnus Ruud
Thapa, Rahul
Pedersen, Magnus Guldberg
Storgaard, Anton Mosquera
Williams, Nick
Gatej, Radu
Lehn-Schiøler, Tue
Brink-Kjær, Andreas
Puthusserypady, Sadasivan
Beniczky, Sándor
Zou, James
Hansen, Lars Kai
author_facet Lehn-Schiøler, William
Kjær, Magnus Ruud
Thapa, Rahul
Pedersen, Magnus Guldberg
Storgaard, Anton Mosquera
Williams, Nick
Gatej, Radu
Lehn-Schiøler, Tue
Brink-Kjær, Andreas
Puthusserypady, Sadasivan
Beniczky, Sándor
Zou, James
Hansen, Lars Kai
contents EEG foundation models achieve state-of-the-art clinical performance, yet the internal computations driving their predictions remain opaque: a barrier to clinical trust. We apply TopK Sparse Autoencoders (SAEs) across three architecturally distinct EEG transformers: SleepFM, REVE, and LaBraM to extract sparse feature dictionaries from their embeddings. By grounding these features in a clinical taxonomy (abnormality, age, sex, and medication), we benchmark monosemanticity and entanglement across architectures. A single hyperparameter procedure, driven by an intrinsic dictionary health audit, transfers robustly across all three architectures. Via concept steering, we introduce a "target vs. off-target" probe area metric to quantify steering selectivity and reveal three operational regimes: selectively steerable, encoded but entangled, and non-encoded. This framework exposes critical representational failures: "wrecking-ball" interventions that collapse global model performance, and clinical entanglements, such as age-pathology confounding, where it is impossible to suppress one concept without corrupting the other. Finally, a spectral decoder maps these interventions back to the amplitude spectrum, translating latent manipulations into physiologically interpretable frequency signatures, such as pathological slow-wave suppression and $α$-band restoration.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13930
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders
Lehn-Schiøler, William
Kjær, Magnus Ruud
Thapa, Rahul
Pedersen, Magnus Guldberg
Storgaard, Anton Mosquera
Williams, Nick
Gatej, Radu
Lehn-Schiøler, Tue
Brink-Kjær, Andreas
Puthusserypady, Sadasivan
Beniczky, Sándor
Zou, James
Hansen, Lars Kai
Machine Learning
Human-Computer Interaction
Neural and Evolutionary Computing
EEG foundation models achieve state-of-the-art clinical performance, yet the internal computations driving their predictions remain opaque: a barrier to clinical trust. We apply TopK Sparse Autoencoders (SAEs) across three architecturally distinct EEG transformers: SleepFM, REVE, and LaBraM to extract sparse feature dictionaries from their embeddings. By grounding these features in a clinical taxonomy (abnormality, age, sex, and medication), we benchmark monosemanticity and entanglement across architectures. A single hyperparameter procedure, driven by an intrinsic dictionary health audit, transfers robustly across all three architectures. Via concept steering, we introduce a "target vs. off-target" probe area metric to quantify steering selectivity and reveal three operational regimes: selectively steerable, encoded but entangled, and non-encoded. This framework exposes critical representational failures: "wrecking-ball" interventions that collapse global model performance, and clinical entanglements, such as age-pathology confounding, where it is impossible to suppress one concept without corrupting the other. Finally, a spectral decoder maps these interventions back to the amplitude spectrum, translating latent manipulations into physiologically interpretable frequency signatures, such as pathological slow-wave suppression and $α$-band restoration.
title Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders
topic Machine Learning
Human-Computer Interaction
Neural and Evolutionary Computing
url https://arxiv.org/abs/2605.13930