EEG-CLIP : Learning EEG representations from natural language descriptions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ndir, Tidiane Camaret, Schirrmeister, Robin Tibor, Ball, Tonio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915415635001344
author Ndir, Tidiane Camaret
Schirrmeister, Robin Tibor
Ball, Tonio
author_facet Ndir, Tidiane Camaret
Schirrmeister, Robin Tibor
Ball, Tonio
contents Deep networks for electroencephalogram (EEG) decoding are often only trained to solve one specific task, such as pathology or age decoding. A more general task-agnostic approach is to train deep networks to match a (clinical) EEG recording to its corresponding textual medical report and vice versa. This approach was pioneered in the computer vision domain matching images and their text captions and subsequently allowed to do successful zero-shot decoding using textual class prompts. In this work, we follow this approach and develop a contrastive learning framework, EEG-CLIP, that aligns the EEG time series and the descriptions of the corresponding clinical text in a shared embedding space. We investigated its potential for versatile EEG decoding, evaluating performance in a range of few-shot and zero-shot settings. Overall, we show that EEG-CLIP manages to non-trivially align text and EEG representations. Our work presents a promising approach to learn general EEG representations, which could enable easier analyses of diverse decoding questions through zero-shot decoding or training task-specific models from fewer training examples. The code for reproducing our results is available at https://github.com/tidiane-camaret/EEGClip
format Preprint
id arxiv_https___arxiv_org_abs_2503_16531
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EEG-CLIP : Learning EEG representations from natural language descriptions
Ndir, Tidiane Camaret
Schirrmeister, Robin Tibor
Ball, Tonio
Computation and Language
Machine Learning
Signal Processing
Deep networks for electroencephalogram (EEG) decoding are often only trained to solve one specific task, such as pathology or age decoding. A more general task-agnostic approach is to train deep networks to match a (clinical) EEG recording to its corresponding textual medical report and vice versa. This approach was pioneered in the computer vision domain matching images and their text captions and subsequently allowed to do successful zero-shot decoding using textual class prompts. In this work, we follow this approach and develop a contrastive learning framework, EEG-CLIP, that aligns the EEG time series and the descriptions of the corresponding clinical text in a shared embedding space. We investigated its potential for versatile EEG decoding, evaluating performance in a range of few-shot and zero-shot settings. Overall, we show that EEG-CLIP manages to non-trivially align text and EEG representations. Our work presents a promising approach to learn general EEG representations, which could enable easier analyses of diverse decoding questions through zero-shot decoding or training task-specific models from fewer training examples. The code for reproducing our results is available at https://github.com/tidiane-camaret/EEGClip
title EEG-CLIP : Learning EEG representations from natural language descriptions
topic Computation and Language
Machine Learning
Signal Processing
url https://arxiv.org/abs/2503.16531