M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fan, Cunhang, Chen, Ying, Zhou, Jian, Pan, Zexu, Zhang, Jingjing, Gao, Youdian, Yang, Xiaoke, Wen, Zhengqi, Lv, Zhao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915315879772160
author Fan, Cunhang
Chen, Ying
Zhou, Jian
Pan, Zexu
Zhang, Jingjing
Gao, Youdian
Yang, Xiaoke
Wen, Zhengqi
Lv, Zhao
author_facet Fan, Cunhang
Chen, Ying
Zhou, Jian
Pan, Zexu
Zhang, Jingjing
Gao, Youdian
Yang, Xiaoke
Wen, Zhengqi
Lv, Zhao
contents The brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlook the issue of temporal misalignment between speech and EEG modalities, which hampers TSE performance. In addition, the speech encoder in current models typically uses basic temporal operations (e.g., one-dimensional convolution), which are unable to effectively extract target speaker information. To address these issues, this paper proposes a multi-scale and multi-modal alignment network (M3ANet) for brain-assisted TSE. Specifically, to eliminate the temporal inconsistency between EEG and speech modalities, the modal alignment module that uses a contrastive learning strategy is applied to align the temporal features of both modalities. Additionally, to fully extract speech information, multi-scale convolutions with GroupMamba modules are used as the speech encoder, which scans speech features at each scale from different directions, enabling the model to capture deep sequence information. Experimental results on three publicly available datasets show that the proposed model outperforms current state-of-the-art methods across various evaluation metrics, highlighting the effectiveness of our proposed method. The source code is available at: https://github.com/fchest/M3ANet.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00466
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction
Fan, Cunhang
Chen, Ying
Zhou, Jian
Pan, Zexu
Zhang, Jingjing
Gao, Youdian
Yang, Xiaoke
Wen, Zhengqi
Lv, Zhao
Audio and Speech Processing
Sound
The brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlook the issue of temporal misalignment between speech and EEG modalities, which hampers TSE performance. In addition, the speech encoder in current models typically uses basic temporal operations (e.g., one-dimensional convolution), which are unable to effectively extract target speaker information. To address these issues, this paper proposes a multi-scale and multi-modal alignment network (M3ANet) for brain-assisted TSE. Specifically, to eliminate the temporal inconsistency between EEG and speech modalities, the modal alignment module that uses a contrastive learning strategy is applied to align the temporal features of both modalities. Additionally, to fully extract speech information, multi-scale convolutions with GroupMamba modules are used as the speech encoder, which scans speech features at each scale from different directions, enabling the model to capture deep sequence information. Experimental results on three publicly available datasets show that the proposed model outperforms current state-of-the-art methods across various evaluation metrics, highlighting the effectiveness of our proposed method. The source code is available at: https://github.com/fchest/M3ANet.
title M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2506.00466