BUT Systems for WildSpoof Challenge: SASV in the Wild

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Junyi, Li, Jin, Rohdin, Johan, Zhang, Lin, Hlaváček, Miroslav, Plchot, Oldrich
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908711598948352
author Peng, Junyi
Li, Jin
Rohdin, Johan
Zhang, Lin
Hlaváček, Miroslav
Plchot, Oldrich
author_facet Peng, Junyi
Li, Jin
Rohdin, Johan
Zhang, Lin
Hlaváček, Miroslav
Plchot, Oldrich
contents This paper presents the BUT submission to the WildSpoof Challenge, focusing on the Spoofing-robust Automatic Speaker Verification (SASV) track. We propose a SASV framework designed to bridge the gap between general audio understanding and specialized speech analysis. Our subsystem integrates diverse Self-Supervised Learning front-ends ranging from general audio models (e.g., Dasheng) to speech-specific encoders (e.g., WavLM). These representations are aggregated via a lightweight Multi-Head Factorized Attention back-end for corresponding subtasks. Furthermore, we introduce a feature domain augmentation strategy based on Distribution Uncertainty to explicitly model and mitigate the domain shift caused by unseen neural vocoders and recording environments. By fusing these robust CM scores with state-of-the-art ASV systems, our approach achieves superior minimization of the a-DCFs and EERs.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12851
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BUT Systems for WildSpoof Challenge: SASV in the Wild
Peng, Junyi
Li, Jin
Rohdin, Johan
Zhang, Lin
Hlaváček, Miroslav
Plchot, Oldrich
Audio and Speech Processing
This paper presents the BUT submission to the WildSpoof Challenge, focusing on the Spoofing-robust Automatic Speaker Verification (SASV) track. We propose a SASV framework designed to bridge the gap between general audio understanding and specialized speech analysis. Our subsystem integrates diverse Self-Supervised Learning front-ends ranging from general audio models (e.g., Dasheng) to speech-specific encoders (e.g., WavLM). These representations are aggregated via a lightweight Multi-Head Factorized Attention back-end for corresponding subtasks. Furthermore, we introduce a feature domain augmentation strategy based on Distribution Uncertainty to explicitly model and mitigate the domain shift caused by unseen neural vocoders and recording environments. By fusing these robust CM scores with state-of-the-art ASV systems, our approach achieves superior minimization of the a-DCFs and EERs.
title BUT Systems for WildSpoof Challenge: SASV in the Wild
topic Audio and Speech Processing
url https://arxiv.org/abs/2512.12851