ATMM-SAGA: Alternating Training for Multi-Module with Score-Aware Gated Attention SASV system

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Asali, Amro, Ben-Shimol, Yehuda, Lapidot, Itshak
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915303155302400
author Asali, Amro
Ben-Shimol, Yehuda
Lapidot, Itshak
author_facet Asali, Amro
Ben-Shimol, Yehuda
Lapidot, Itshak
contents The objective of automatic speaker verification (ASV) systems is to determine whether a given test speech utterance corresponds to a claimed enrolled speaker. These systems have a wide range of applications, and ensuring their reliability is crucial. In this paper, we propose a spoofing-robust automatic speaker verification (SASV) system employing a score-aware gated attention (SAGA) fusion scheme, integrating scores from a pre-trained countermeasure (CM) with speaker embeddings from a pre-trained ASV. Specifically, we employ the AASIST and ECAPA-TDNN models. SAGA acts as an adaptive gating mechanism, where the CM score determines how strongly ASV embeddings influence the final SASV decision. Experiments on the ASVspoof2019 logical access dataset demonstrate that the proposed SASV system achieves an SASV equal error rate (SASV-EER) and agnostic detection cost function (a-DCF) of 2.31%, 0.0603 for the development set and 2.18%, 0.0480 for the evaluation set.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18273
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ATMM-SAGA: Alternating Training for Multi-Module with Score-Aware Gated Attention SASV system
Asali, Amro
Ben-Shimol, Yehuda
Lapidot, Itshak
Audio and Speech Processing
The objective of automatic speaker verification (ASV) systems is to determine whether a given test speech utterance corresponds to a claimed enrolled speaker. These systems have a wide range of applications, and ensuring their reliability is crucial. In this paper, we propose a spoofing-robust automatic speaker verification (SASV) system employing a score-aware gated attention (SAGA) fusion scheme, integrating scores from a pre-trained countermeasure (CM) with speaker embeddings from a pre-trained ASV. Specifically, we employ the AASIST and ECAPA-TDNN models. SAGA acts as an adaptive gating mechanism, where the CM score determines how strongly ASV embeddings influence the final SASV decision. Experiments on the ASVspoof2019 logical access dataset demonstrate that the proposed SASV system achieves an SASV equal error rate (SASV-EER) and agnostic detection cost function (a-DCF) of 2.31%, 0.0603 for the development set and 2.18%, 0.0480 for the evaluation set.
title ATMM-SAGA: Alternating Training for Multi-Module with Score-Aware Gated Attention SASV system
topic Audio and Speech Processing
url https://arxiv.org/abs/2505.18273