Spoofing-Aware Speaker Verification via Wavelet Prompt Tuning and Multi-Model Ensembles

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Farhadipour, Aref, Jin, Ming, Vyshnevetska, Valeriia, Li, Xiyang, Pellegrino, Elisa, Madikeri, Srikanth
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917221084692480
author Farhadipour, Aref
Jin, Ming
Vyshnevetska, Valeriia
Li, Xiyang
Pellegrino, Elisa
Madikeri, Srikanth
author_facet Farhadipour, Aref
Jin, Ming
Vyshnevetska, Valeriia
Li, Xiyang
Pellegrino, Elisa
Madikeri, Srikanth
contents This paper describes the UZH-CL system submitted to the SASV section of the WildSpoof 2026 challenge. The challenge focuses on the integrated defense against generative spoofing attacks by requiring the simultaneous verification of speaker identity and audio authenticity. We proposed a cascaded Spoofing-Aware Speaker Verification framework that integrates a Wavelet Prompt-Tuned XLSR-AASIST countermeasure with a multi-model ensemble. The ASV component utilizes the ResNet34, ResNet293, and WavLM-ECAPA-TDNN architectures, with Z-score normalization followed by score averaging. Trained on VoxCeleb2 and SpoofCeleb, the system obtained a Macro a-DCF of 0.2017 and a SASV EER of 2.08%. While the system achieved a 0.16% EER in spoof detection on the in-domain data, results on unseen datasets, such as the ASVspoof5, highlight the critical challenge of cross-domain generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17557
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Spoofing-Aware Speaker Verification via Wavelet Prompt Tuning and Multi-Model Ensembles
Farhadipour, Aref
Jin, Ming
Vyshnevetska, Valeriia
Li, Xiyang
Pellegrino, Elisa
Madikeri, Srikanth
Audio and Speech Processing
Sound
This paper describes the UZH-CL system submitted to the SASV section of the WildSpoof 2026 challenge. The challenge focuses on the integrated defense against generative spoofing attacks by requiring the simultaneous verification of speaker identity and audio authenticity. We proposed a cascaded Spoofing-Aware Speaker Verification framework that integrates a Wavelet Prompt-Tuned XLSR-AASIST countermeasure with a multi-model ensemble. The ASV component utilizes the ResNet34, ResNet293, and WavLM-ECAPA-TDNN architectures, with Z-score normalization followed by score averaging. Trained on VoxCeleb2 and SpoofCeleb, the system obtained a Macro a-DCF of 0.2017 and a SASV EER of 2.08%. While the system achieved a 0.16% EER in spoof detection on the in-domain data, results on unseen datasets, such as the ASVspoof5, highlight the critical challenge of cross-domain generalization.
title Spoofing-Aware Speaker Verification via Wavelet Prompt Tuning and Multi-Model Ensembles
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2601.17557