Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917524463943680 |
|---|---|
| author | Lehn-Schiøler, William Kjær, Magnus Ruud Thapa, Rahul Pedersen, Magnus Guldberg Storgaard, Anton Mosquera Williams, Nick Gatej, Radu Lehn-Schiøler, Tue Brink-Kjær, Andreas Puthusserypady, Sadasivan Beniczky, Sándor Zou, James Hansen, Lars Kai |
| author_facet | Lehn-Schiøler, William Kjær, Magnus Ruud Thapa, Rahul Pedersen, Magnus Guldberg Storgaard, Anton Mosquera Williams, Nick Gatej, Radu Lehn-Schiøler, Tue Brink-Kjær, Andreas Puthusserypady, Sadasivan Beniczky, Sándor Zou, James Hansen, Lars Kai |
| contents | EEG foundation models achieve state-of-the-art clinical performance, yet the internal computations driving their predictions remain opaque: a barrier to clinical trust. We apply TopK Sparse Autoencoders (SAEs) across three architecturally distinct EEG transformers: SleepFM, REVE, and LaBraM to extract sparse feature dictionaries from their embeddings. By grounding these features in a clinical taxonomy (abnormality, age, sex, and medication), we benchmark monosemanticity and entanglement across architectures. A single hyperparameter procedure, driven by an intrinsic dictionary health audit, transfers robustly across all three architectures. Via concept steering, we introduce a "target vs. off-target" probe area metric to quantify steering selectivity and reveal three operational regimes: selectively steerable, encoded but entangled, and non-encoded. This framework exposes critical representational failures: "wrecking-ball" interventions that collapse global model performance, and clinical entanglements, such as age-pathology confounding, where it is impossible to suppress one concept without corrupting the other. Finally, a spectral decoder maps these interventions back to the amplitude spectrum, translating latent manipulations into physiologically interpretable frequency signatures, such as pathological slow-wave suppression and $α$-band restoration. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_13930 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders Lehn-Schiøler, William Kjær, Magnus Ruud Thapa, Rahul Pedersen, Magnus Guldberg Storgaard, Anton Mosquera Williams, Nick Gatej, Radu Lehn-Schiøler, Tue Brink-Kjær, Andreas Puthusserypady, Sadasivan Beniczky, Sándor Zou, James Hansen, Lars Kai Machine Learning Human-Computer Interaction Neural and Evolutionary Computing EEG foundation models achieve state-of-the-art clinical performance, yet the internal computations driving their predictions remain opaque: a barrier to clinical trust. We apply TopK Sparse Autoencoders (SAEs) across three architecturally distinct EEG transformers: SleepFM, REVE, and LaBraM to extract sparse feature dictionaries from their embeddings. By grounding these features in a clinical taxonomy (abnormality, age, sex, and medication), we benchmark monosemanticity and entanglement across architectures. A single hyperparameter procedure, driven by an intrinsic dictionary health audit, transfers robustly across all three architectures. Via concept steering, we introduce a "target vs. off-target" probe area metric to quantify steering selectivity and reveal three operational regimes: selectively steerable, encoded but entangled, and non-encoded. This framework exposes critical representational failures: "wrecking-ball" interventions that collapse global model performance, and clinical entanglements, such as age-pathology confounding, where it is impossible to suppress one concept without corrupting the other. Finally, a spectral decoder maps these interventions back to the amplitude spectrum, translating latent manipulations into physiologically interpretable frequency signatures, such as pathological slow-wave suppression and $α$-band restoration. |
| title | Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders |
| topic | Machine Learning Human-Computer Interaction Neural and Evolutionary Computing |
| url | https://arxiv.org/abs/2605.13930 |