RawBMamba: End-to-End Bidirectional State Space Model for Audio Deepfake Detection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Yujie, Yi, Jiangyan, Xue, Jun, Wang, Chenglong, Zhang, Xiaohui, Dong, Shunbo, Zeng, Siding, Tao, Jianhua, Zhao, Lv, Fan, Cunhang
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911923839172608
author Chen, Yujie
Yi, Jiangyan
Xue, Jun
Wang, Chenglong
Zhang, Xiaohui
Dong, Shunbo
Zeng, Siding
Tao, Jianhua
Zhao, Lv
Fan, Cunhang
author_facet Chen, Yujie
Yi, Jiangyan
Xue, Jun
Wang, Chenglong
Zhang, Xiaohui
Dong, Shunbo
Zeng, Siding
Tao, Jianhua
Zhao, Lv
Fan, Cunhang
contents Fake artefacts for discriminating between bonafide and fake audio can exist in both short- and long-range segments. Therefore, combining local and global feature information can effectively discriminate between bonafide and fake audio. This paper proposes an end-to-end bidirectional state space model, named RawBMamba, to capture both short- and long-range discriminative information for audio deepfake detection. Specifically, we use sinc Layer and multiple convolutional layers to capture short-range features, and then design a bidirectional Mamba to address Mamba's unidirectional modelling problem and further capture long-range feature information. Moreover, we develop a bidirectional fusion module to integrate embeddings, enhancing audio context representation and combining short- and long-range information. The results show that our proposed RawBMamba achieves a 34.1\% improvement over Rawformer on ASVspoof2021 LA dataset, and demonstrates competitive performance on other datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2406_06086
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RawBMamba: End-to-End Bidirectional State Space Model for Audio Deepfake Detection
Chen, Yujie
Yi, Jiangyan
Xue, Jun
Wang, Chenglong
Zhang, Xiaohui
Dong, Shunbo
Zeng, Siding
Tao, Jianhua
Zhao, Lv
Fan, Cunhang
Sound
Audio and Speech Processing
Fake artefacts for discriminating between bonafide and fake audio can exist in both short- and long-range segments. Therefore, combining local and global feature information can effectively discriminate between bonafide and fake audio. This paper proposes an end-to-end bidirectional state space model, named RawBMamba, to capture both short- and long-range discriminative information for audio deepfake detection. Specifically, we use sinc Layer and multiple convolutional layers to capture short-range features, and then design a bidirectional Mamba to address Mamba's unidirectional modelling problem and further capture long-range feature information. Moreover, we develop a bidirectional fusion module to integrate embeddings, enhancing audio context representation and combining short- and long-range information. The results show that our proposed RawBMamba achieves a 34.1\% improvement over Rawformer on ASVspoof2021 LA dataset, and demonstrates competitive performance on other datasets.
title RawBMamba: End-to-End Bidirectional State Space Model for Audio Deepfake Detection
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2406.06086