XMUspeech Systems for the ASVspoof 5 Challenge

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Wangjie, Xie, Xingjia, Li, Yishuang, Guan, Wenhao, Wang, Kaidi, Ren, Pengyu, Li, Lin, Hong, Qingyang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915506862161920
author Li, Wangjie
Xie, Xingjia
Li, Yishuang
Guan, Wenhao
Wang, Kaidi
Ren, Pengyu
Li, Lin
Hong, Qingyang
author_facet Li, Wangjie
Xie, Xingjia
Li, Yishuang
Guan, Wenhao
Wang, Kaidi
Ren, Pengyu
Li, Lin
Hong, Qingyang
contents In this paper, we present our submitted XMUspeech systems to the speech deepfake detection track of the ASVspoof 5 Challenge. Compared to previous challenges, the audio duration in ASVspoof 5 database has significantly increased. And we observed that merely adjusting the input audio length can substantially improve system performance. To capture artifacts at multiple levels, we explored the performance of AASIST, HM-Conformer, Hubert, and Wav2vec2 with various input features and loss functions. Specifically, in order to obtain artifact-related information, we trained self-supervised models on the dataset containing spoofing utterances as the feature extractors. And we applied an adaptive multi-scale feature fusion (AMFF) method to integrate features from multiple Transformer layers with the hand-crafted feature to enhance the detection capability. In addition, we conducted extensive experiments on one-class loss functions and provided optimized configurations to better align with the anti-spoofing task. Our fusion system achieved a minDCF of 0.4783 and an EER of 20.45% in the closed condition, and a minDCF of 0.2245 and an EER of 9.36% in the open condition.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18102
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle XMUspeech Systems for the ASVspoof 5 Challenge
Li, Wangjie
Xie, Xingjia
Li, Yishuang
Guan, Wenhao
Wang, Kaidi
Ren, Pengyu
Li, Lin
Hong, Qingyang
Sound
Audio and Speech Processing
In this paper, we present our submitted XMUspeech systems to the speech deepfake detection track of the ASVspoof 5 Challenge. Compared to previous challenges, the audio duration in ASVspoof 5 database has significantly increased. And we observed that merely adjusting the input audio length can substantially improve system performance. To capture artifacts at multiple levels, we explored the performance of AASIST, HM-Conformer, Hubert, and Wav2vec2 with various input features and loss functions. Specifically, in order to obtain artifact-related information, we trained self-supervised models on the dataset containing spoofing utterances as the feature extractors. And we applied an adaptive multi-scale feature fusion (AMFF) method to integrate features from multiple Transformer layers with the hand-crafted feature to enhance the detection capability. In addition, we conducted extensive experiments on one-class loss functions and provided optimized configurations to better align with the anti-spoofing task. Our fusion system achieved a minDCF of 0.4783 and an EER of 20.45% in the closed condition, and a minDCF of 0.2245 and an EER of 9.36% in the open condition.
title XMUspeech Systems for the ASVspoof 5 Challenge
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.18102