Decoding Covert Speech from EEG Using a Functional Areas Spatio-Temporal Transformer

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Jiang, Muyun, Ding, Yi, Zhang, Wei, Teo, Kok Ann Colin, Fong, LaiGuan, Zhang, Shuailei, Guo, Zhiwei, Liu, Chenyu, Bhuvanakantham, Raghavan, Sim, Wei Khang Jeremy, Foo, Chuan Huat Vince, Chua, Rong Hui Jonathan, Padmanabhan, Parasuraman, Leong, Victoria, Lu, Jia, Gulyas, Balazs, Guan, Cuntai
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910904157732864
author Jiang, Muyun
Ding, Yi
Zhang, Wei
Teo, Kok Ann Colin
Fong, LaiGuan
Zhang, Shuailei
Guo, Zhiwei
Liu, Chenyu
Bhuvanakantham, Raghavan
Sim, Wei Khang Jeremy
Foo, Chuan Huat Vince
Chua, Rong Hui Jonathan
Padmanabhan, Parasuraman
Leong, Victoria
Lu, Jia
Gulyas, Balazs
Guan, Cuntai
author_facet Jiang, Muyun
Ding, Yi
Zhang, Wei
Teo, Kok Ann Colin
Fong, LaiGuan
Zhang, Shuailei
Guo, Zhiwei
Liu, Chenyu
Bhuvanakantham, Raghavan
Sim, Wei Khang Jeremy
Foo, Chuan Huat Vince
Chua, Rong Hui Jonathan
Padmanabhan, Parasuraman
Leong, Victoria
Lu, Jia
Gulyas, Balazs
Guan, Cuntai
contents Covert speech involves imagining speaking without audible sound or any movements. Decoding covert speech from electroencephalogram (EEG) is challenging due to a limited understanding of neural pronunciation mapping and the low signal-to-noise ratio of the signal. In this study, we developed a large-scale multi-utterance speech EEG dataset from 57 right-handed native English-speaking subjects, each performing covert and overt speech tasks by repeating the same word in five utterances within a ten-second duration. Given the spatio-temporal nature of the neural activation process during speech pronunciation, we developed a Functional Areas Spatio-temporal Transformer (FAST), an effective framework for converting EEG signals into tokens and utilizing transformer architecture for sequence encoding. Our results reveal distinct and interpretable speech neural features by the visualization of FAST-generated activation maps across frontal and temporal brain regions with each word being covertly spoken, providing new insights into the discriminative features of the neural representation of covert speech. This is the first report of such a study, which provides interpretable evidence for speech decoding from EEG. The code for this work has been made public at https://github.com/Jiang-Muyun/FAST
format Preprint
id arxiv_https___arxiv_org_abs_2504_03762
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Decoding Covert Speech from EEG Using a Functional Areas Spatio-Temporal Transformer
Jiang, Muyun
Ding, Yi
Zhang, Wei
Teo, Kok Ann Colin
Fong, LaiGuan
Zhang, Shuailei
Guo, Zhiwei
Liu, Chenyu
Bhuvanakantham, Raghavan
Sim, Wei Khang Jeremy
Foo, Chuan Huat Vince
Chua, Rong Hui Jonathan
Padmanabhan, Parasuraman
Leong, Victoria
Lu, Jia
Gulyas, Balazs
Guan, Cuntai
Signal Processing
Machine Learning
Covert speech involves imagining speaking without audible sound or any movements. Decoding covert speech from electroencephalogram (EEG) is challenging due to a limited understanding of neural pronunciation mapping and the low signal-to-noise ratio of the signal. In this study, we developed a large-scale multi-utterance speech EEG dataset from 57 right-handed native English-speaking subjects, each performing covert and overt speech tasks by repeating the same word in five utterances within a ten-second duration. Given the spatio-temporal nature of the neural activation process during speech pronunciation, we developed a Functional Areas Spatio-temporal Transformer (FAST), an effective framework for converting EEG signals into tokens and utilizing transformer architecture for sequence encoding. Our results reveal distinct and interpretable speech neural features by the visualization of FAST-generated activation maps across frontal and temporal brain regions with each word being covertly spoken, providing new insights into the discriminative features of the neural representation of covert speech. This is the first report of such a study, which provides interpretable evidence for speech decoding from EEG. The code for this work has been made public at https://github.com/Jiang-Muyun/FAST
title Decoding Covert Speech from EEG Using a Functional Areas Spatio-Temporal Transformer
topic Signal Processing
Machine Learning
url https://arxiv.org/abs/2504.03762