Rapidly Adapting to New Voice Spoofing: Few-Shot Detection of Synthesized Speech Under Distribution Shifts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Garg, Ashi, Cai, Zexin, Xinyuan, Henry Li, García-Perera, Leibny Paola, Duh, Kevin, Khudanpur, Sanjeev, Wiesner, Matthew, Andrews, Nicholas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913996553060352
author Garg, Ashi
Cai, Zexin
Xinyuan, Henry Li
García-Perera, Leibny Paola
Duh, Kevin
Khudanpur, Sanjeev
Wiesner, Matthew
Andrews, Nicholas
author_facet Garg, Ashi
Cai, Zexin
Xinyuan, Henry Li
García-Perera, Leibny Paola
Duh, Kevin
Khudanpur, Sanjeev
Wiesner, Matthew
Andrews, Nicholas
contents We address the challenge of detecting synthesized speech under distribution shifts -- arising from unseen synthesis methods, speakers, languages, or audio conditions -- relative to the training data. Few-shot learning methods are a promising way to tackle distribution shifts by rapidly adapting on the basis of a few in-distribution samples. We propose a self-attentive prototypical network to enable more robust few-shot adaptation. To evaluate our approach, we systematically compare the performance of traditional zero-shot detectors and the proposed few-shot detectors, carefully controlling training conditions to introduce distribution shifts at evaluation time. In conditions where distribution shifts hamper the zero-shot performance, our proposed few-shot adaptation technique can quickly adapt using as few as 10 in-distribution samples -- achieving upto 32% relative EER reduction on deepfakes in Japanese language and 20% relative reduction on ASVspoof 2021 Deepfake dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2508_13320
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rapidly Adapting to New Voice Spoofing: Few-Shot Detection of Synthesized Speech Under Distribution Shifts
Garg, Ashi
Cai, Zexin
Xinyuan, Henry Li
García-Perera, Leibny Paola
Duh, Kevin
Khudanpur, Sanjeev
Wiesner, Matthew
Andrews, Nicholas
Audio and Speech Processing
We address the challenge of detecting synthesized speech under distribution shifts -- arising from unseen synthesis methods, speakers, languages, or audio conditions -- relative to the training data. Few-shot learning methods are a promising way to tackle distribution shifts by rapidly adapting on the basis of a few in-distribution samples. We propose a self-attentive prototypical network to enable more robust few-shot adaptation. To evaluate our approach, we systematically compare the performance of traditional zero-shot detectors and the proposed few-shot detectors, carefully controlling training conditions to introduce distribution shifts at evaluation time. In conditions where distribution shifts hamper the zero-shot performance, our proposed few-shot adaptation technique can quickly adapt using as few as 10 in-distribution samples -- achieving upto 32% relative EER reduction on deepfakes in Japanese language and 20% relative reduction on ASVspoof 2021 Deepfake dataset.
title Rapidly Adapting to New Voice Spoofing: Few-Shot Detection of Synthesized Speech Under Distribution Shifts
topic Audio and Speech Processing
url https://arxiv.org/abs/2508.13320