XMUspeech Systems for the ASVspoof 5 Challenge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915506862161920 |
|---|---|
| author | Li, Wangjie Xie, Xingjia Li, Yishuang Guan, Wenhao Wang, Kaidi Ren, Pengyu Li, Lin Hong, Qingyang |
| author_facet | Li, Wangjie Xie, Xingjia Li, Yishuang Guan, Wenhao Wang, Kaidi Ren, Pengyu Li, Lin Hong, Qingyang |
| contents | In this paper, we present our submitted XMUspeech systems to the speech deepfake detection track of the ASVspoof 5 Challenge. Compared to previous challenges, the audio duration in ASVspoof 5 database has significantly increased. And we observed that merely adjusting the input audio length can substantially improve system performance. To capture artifacts at multiple levels, we explored the performance of AASIST, HM-Conformer, Hubert, and Wav2vec2 with various input features and loss functions. Specifically, in order to obtain artifact-related information, we trained self-supervised models on the dataset containing spoofing utterances as the feature extractors. And we applied an adaptive multi-scale feature fusion (AMFF) method to integrate features from multiple Transformer layers with the hand-crafted feature to enhance the detection capability. In addition, we conducted extensive experiments on one-class loss functions and provided optimized configurations to better align with the anti-spoofing task. Our fusion system achieved a minDCF of 0.4783 and an EER of 20.45% in the closed condition, and a minDCF of 0.2245 and an EER of 9.36% in the open condition. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_18102 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | XMUspeech Systems for the ASVspoof 5 Challenge Li, Wangjie Xie, Xingjia Li, Yishuang Guan, Wenhao Wang, Kaidi Ren, Pengyu Li, Lin Hong, Qingyang Sound Audio and Speech Processing In this paper, we present our submitted XMUspeech systems to the speech deepfake detection track of the ASVspoof 5 Challenge. Compared to previous challenges, the audio duration in ASVspoof 5 database has significantly increased. And we observed that merely adjusting the input audio length can substantially improve system performance. To capture artifacts at multiple levels, we explored the performance of AASIST, HM-Conformer, Hubert, and Wav2vec2 with various input features and loss functions. Specifically, in order to obtain artifact-related information, we trained self-supervised models on the dataset containing spoofing utterances as the feature extractors. And we applied an adaptive multi-scale feature fusion (AMFF) method to integrate features from multiple Transformer layers with the hand-crafted feature to enhance the detection capability. In addition, we conducted extensive experiments on one-class loss functions and provided optimized configurations to better align with the anti-spoofing task. Our fusion system achieved a minDCF of 0.4783 and an EER of 20.45% in the closed condition, and a minDCF of 0.2245 and an EER of 9.36% in the open condition. |
| title | XMUspeech Systems for the ASVspoof 5 Challenge |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2509.18102 |