BUT Systems for WildSpoof Challenge: SASV in the Wild
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908711598948352 |
|---|---|
| author | Peng, Junyi Li, Jin Rohdin, Johan Zhang, Lin Hlaváček, Miroslav Plchot, Oldrich |
| author_facet | Peng, Junyi Li, Jin Rohdin, Johan Zhang, Lin Hlaváček, Miroslav Plchot, Oldrich |
| contents | This paper presents the BUT submission to the WildSpoof Challenge, focusing on the Spoofing-robust Automatic Speaker Verification (SASV) track. We propose a SASV framework designed to bridge the gap between general audio understanding and specialized speech analysis. Our subsystem integrates diverse Self-Supervised Learning front-ends ranging from general audio models (e.g., Dasheng) to speech-specific encoders (e.g., WavLM). These representations are aggregated via a lightweight Multi-Head Factorized Attention back-end for corresponding subtasks. Furthermore, we introduce a feature domain augmentation strategy based on Distribution Uncertainty to explicitly model and mitigate the domain shift caused by unseen neural vocoders and recording environments. By fusing these robust CM scores with state-of-the-art ASV systems, our approach achieves superior minimization of the a-DCFs and EERs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_12851 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | BUT Systems for WildSpoof Challenge: SASV in the Wild Peng, Junyi Li, Jin Rohdin, Johan Zhang, Lin Hlaváček, Miroslav Plchot, Oldrich Audio and Speech Processing This paper presents the BUT submission to the WildSpoof Challenge, focusing on the Spoofing-robust Automatic Speaker Verification (SASV) track. We propose a SASV framework designed to bridge the gap between general audio understanding and specialized speech analysis. Our subsystem integrates diverse Self-Supervised Learning front-ends ranging from general audio models (e.g., Dasheng) to speech-specific encoders (e.g., WavLM). These representations are aggregated via a lightweight Multi-Head Factorized Attention back-end for corresponding subtasks. Furthermore, we introduce a feature domain augmentation strategy based on Distribution Uncertainty to explicitly model and mitigate the domain shift caused by unseen neural vocoders and recording environments. By fusing these robust CM scores with state-of-the-art ASV systems, our approach achieves superior minimization of the a-DCFs and EERs. |
| title | BUT Systems for WildSpoof Challenge: SASV in the Wild |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2512.12851 |