SPGM: Prioritizing Local Features for enhanced speech separation performance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yip, Jia Qi, Zhao, Shengkui, Ma, Yukun, Ni, Chongjia, Zhang, Chong, Wang, Hao, Nguyen, Trung Hieu, Zhou, Kun, Ng, Dianwen, Chng, Eng Siong, Ma, Bin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929269201960960
author Yip, Jia Qi
Zhao, Shengkui
Ma, Yukun
Ni, Chongjia
Zhang, Chong
Wang, Hao
Nguyen, Trung Hieu
Zhou, Kun
Ng, Dianwen
Chng, Eng Siong
Ma, Bin
author_facet Yip, Jia Qi
Zhao, Shengkui
Ma, Yukun
Ni, Chongjia
Zhang, Chong
Wang, Hao
Nguyen, Trung Hieu
Zhou, Kun
Ng, Dianwen
Chng, Eng Siong
Ma, Bin
contents Dual-path is a popular architecture for speech separation models (e.g. Sepformer) which splits long sequences into overlapping chunks for its intra- and inter-blocks that separately model intra-chunk local features and inter-chunk global relationships. However, it has been found that inter-blocks, which comprise half a dual-path model's parameters, contribute minimally to performance. Thus, we propose the Single-Path Global Modulation (SPGM) block to replace inter-blocks. SPGM is named after its structure consisting of a parameter-free global pooling module followed by a modulation module comprising only 2% of the model's total parameters. The SPGM block allows all transformer layers in the model to be dedicated to local feature modelling, making the overall model single-path. SPGM achieves 22.1 dB SI-SDRi on WSJ0-2Mix and 20.4 dB SI-SDRi on Libri2Mix, exceeding the performance of Sepformer by 0.5 dB and 0.3 dB respectively and matches the performance of recent SOTA models with up to 8 times fewer parameters. Model and weights are available at huggingface.co/yipjiaqi/spgm
format Preprint
id arxiv_https___arxiv_org_abs_2309_12608
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SPGM: Prioritizing Local Features for enhanced speech separation performance
Yip, Jia Qi
Zhao, Shengkui
Ma, Yukun
Ni, Chongjia
Zhang, Chong
Wang, Hao
Nguyen, Trung Hieu
Zhou, Kun
Ng, Dianwen
Chng, Eng Siong
Ma, Bin
Audio and Speech Processing
Sound
Dual-path is a popular architecture for speech separation models (e.g. Sepformer) which splits long sequences into overlapping chunks for its intra- and inter-blocks that separately model intra-chunk local features and inter-chunk global relationships. However, it has been found that inter-blocks, which comprise half a dual-path model's parameters, contribute minimally to performance. Thus, we propose the Single-Path Global Modulation (SPGM) block to replace inter-blocks. SPGM is named after its structure consisting of a parameter-free global pooling module followed by a modulation module comprising only 2% of the model's total parameters. The SPGM block allows all transformer layers in the model to be dedicated to local feature modelling, making the overall model single-path. SPGM achieves 22.1 dB SI-SDRi on WSJ0-2Mix and 20.4 dB SI-SDRi on Libri2Mix, exceeding the performance of Sepformer by 0.5 dB and 0.3 dB respectively and matches the performance of recent SOTA models with up to 8 times fewer parameters. Model and weights are available at huggingface.co/yipjiaqi/spgm
title SPGM: Prioritizing Local Features for enhanced speech separation performance
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2309.12608