XLSR-Mamba: A Dual-Column Bidirectional State Space Model for Spoofing Attack Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Yang, Das, Rohan Kumar
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915178200694784
author Xiao, Yang
Das, Rohan Kumar
author_facet Xiao, Yang
Das, Rohan Kumar
contents Transformers and their variants have achieved great success in speech processing. However, their multi-head self-attention mechanism is computationally expensive. Therefore, one novel selective state space model, Mamba, has been proposed as an alternative. Building on its success in automatic speech recognition, we apply Mamba for spoofing attack detection. Mamba is well-suited for this task as it can capture the artifacts in spoofed speech signals by handling long-length sequences. However, Mamba's performance may suffer when it is trained with limited labeled data. To mitigate this, we propose combining a new structure of Mamba based on a dual-column architecture with self-supervised learning, using the pre-trained wav2vec 2.0 model. The experiments show that our proposed approach achieves competitive results and faster inference on the ASVspoof 2021 LA and DF datasets, and on the more challenging In-the-Wild dataset, it emerges as the strongest candidate for spoofing attack detection. The code has been publicly released in https://github.com/swagshaw/XLSR-Mamba.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10027
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle XLSR-Mamba: A Dual-Column Bidirectional State Space Model for Spoofing Attack Detection
Xiao, Yang
Das, Rohan Kumar
Audio and Speech Processing
Sound
Transformers and their variants have achieved great success in speech processing. However, their multi-head self-attention mechanism is computationally expensive. Therefore, one novel selective state space model, Mamba, has been proposed as an alternative. Building on its success in automatic speech recognition, we apply Mamba for spoofing attack detection. Mamba is well-suited for this task as it can capture the artifacts in spoofed speech signals by handling long-length sequences. However, Mamba's performance may suffer when it is trained with limited labeled data. To mitigate this, we propose combining a new structure of Mamba based on a dual-column architecture with self-supervised learning, using the pre-trained wav2vec 2.0 model. The experiments show that our proposed approach achieves competitive results and faster inference on the ASVspoof 2021 LA and DF datasets, and on the more challenging In-the-Wild dataset, it emerges as the strongest candidate for spoofing attack detection. The code has been publicly released in https://github.com/swagshaw/XLSR-Mamba.
title XLSR-Mamba: A Dual-Column Bidirectional State Space Model for Spoofing Attack Detection
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2411.10027