Self-supervised speech representation and contextual text embedding for match-mismatch classification with EEG recording

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Bo, Xu, Xiran, Zhang, Zechen, Zhu, Haolin, Yan, YuJie, Wu, Xihong, Chen, Jing
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916111626272768
author Wang, Bo
Xu, Xiran
Zhang, Zechen
Zhu, Haolin
Yan, YuJie
Wu, Xihong
Chen, Jing
author_facet Wang, Bo
Xu, Xiran
Zhang, Zechen
Zhu, Haolin
Yan, YuJie
Wu, Xihong
Chen, Jing
contents Relating speech to EEG holds considerable importance but is challenging. In this study, a deep convolutional network was employed to extract spatiotemporal features from EEG data. Self-supervised speech representation and contextual text embedding were used as speech features. Contrastive learning was used to relate EEG features to speech features. The experimental results demonstrate the benefits of using self-supervised speech representation and contextual text embedding. Through feature fusion and model ensemble, an accuracy of 60.29% was achieved, and the performance was ranked as No.2 in Task 1 of the Auditory EEG Challenge (ICASSP 2024). The code to implement our work is available on Github: https://github.com/bobwangPKU/EEG-Stimulus-Match-Mismatch.
format Preprint
id arxiv_https___arxiv_org_abs_2401_04964
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Self-supervised speech representation and contextual text embedding for match-mismatch classification with EEG recording
Wang, Bo
Xu, Xiran
Zhang, Zechen
Zhu, Haolin
Yan, YuJie
Wu, Xihong
Chen, Jing
Signal Processing
Sound
Audio and Speech Processing
Relating speech to EEG holds considerable importance but is challenging. In this study, a deep convolutional network was employed to extract spatiotemporal features from EEG data. Self-supervised speech representation and contextual text embedding were used as speech features. Contrastive learning was used to relate EEG features to speech features. The experimental results demonstrate the benefits of using self-supervised speech representation and contextual text embedding. Through feature fusion and model ensemble, an accuracy of 60.29% was achieved, and the performance was ranked as No.2 in Task 1 of the Auditory EEG Challenge (ICASSP 2024). The code to implement our work is available on Github: https://github.com/bobwangPKU/EEG-Stimulus-Match-Mismatch.
title Self-supervised speech representation and contextual text embedding for match-mismatch classification with EEG recording
topic Signal Processing
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2401.04964