SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fan, Cunhang, Zhang, Sheng, Zhang, Jingjing, Pan, Zexu, Lv, Zhao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909460365049856
author Fan, Cunhang
Zhang, Sheng
Zhang, Jingjing
Pan, Zexu
Lv, Zhao
author_facet Fan, Cunhang
Zhang, Sheng
Zhang, Jingjing
Pan, Zexu
Lv, Zhao
contents Decoding speech from brain signals is a challenging research problem that holds significant importance for studying speech processing in the brain. Although breakthroughs have been made in reconstructing the mel spectrograms of audio stimuli perceived by subjects at the word or letter level using noninvasive electroencephalography (EEG), there is still a critical gap in precisely reconstructing continuous speech features, especially at the minute level. To address this issue, this paper proposes a State Space Model (SSM) to reconstruct the mel spectrogram of continuous speech from EEG, named SSM2Mel. This model introduces a novel Mamba module to effectively model the long sequence of EEG signals for imagined speech. In the SSM2Mel model, the S4-UNet structure is used to enhance the extraction of local features of EEG signals, and the Embedding Strength Modulator (ESM) module is used to incorporate subject-specific information. Experimental results show that our model achieves a Pearson correlation of 0.069 on the SparrKULee dataset, which is a 38% improvement over the previous baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2501_10402
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG
Fan, Cunhang
Zhang, Sheng
Zhang, Jingjing
Pan, Zexu
Lv, Zhao
Signal Processing
Sound
Audio and Speech Processing
Decoding speech from brain signals is a challenging research problem that holds significant importance for studying speech processing in the brain. Although breakthroughs have been made in reconstructing the mel spectrograms of audio stimuli perceived by subjects at the word or letter level using noninvasive electroencephalography (EEG), there is still a critical gap in precisely reconstructing continuous speech features, especially at the minute level. To address this issue, this paper proposes a State Space Model (SSM) to reconstruct the mel spectrogram of continuous speech from EEG, named SSM2Mel. This model introduces a novel Mamba module to effectively model the long sequence of EEG signals for imagined speech. In the SSM2Mel model, the S4-UNet structure is used to enhance the extraction of local features of EEG signals, and the Embedding Strength Modulator (ESM) module is used to incorporate subject-specific information. Experimental results show that our model achieves a Pearson correlation of 0.069 on the SparrKULee dataset, which is a 38% improvement over the previous baseline.
title SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG
topic Signal Processing
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2501.10402