Why Speech Deepfake Detectors Won't Generalize: The Limits of Detection in an Open World

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Berisha, Visar, Kadambi, Prad, Lenz, Isabella
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908557015777280
author Berisha, Visar
Kadambi, Prad
Lenz, Isabella
author_facet Berisha, Visar
Kadambi, Prad
Lenz, Isabella
contents Speech deepfake detectors are often evaluated on clean, benchmark-style conditions, but deployment occurs in an open world of shifting devices, sampling rates, codecs, environments, and attack families. This creates a ``coverage debt" for AI-based detectors: every new condition multiplies with existing ones, producing data blind spots that grow faster than data can be collected. Because attackers can target these uncovered regions, worst-case performance (not average benchmark scores) determines security. To demonstrate the impact of the coverage debt problem, we analyze results from a recent cross-testing framework. Grouping performance by bona fide domain and spoof release year, two patterns emerge: newer synthesizers erase the legacy artifacts detectors rely on, and conversational speech domains (teleconferencing, interviews, social media) are consistently the hardest to secure. These findings show that detection alone should not be relied upon for high-stakes decisions. Detectors should be treated as auxiliary signals within layered defenses that include provenance, personhood credentials, and policy safeguards.
format Preprint
id arxiv_https___arxiv_org_abs_2509_20405
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Why Speech Deepfake Detectors Won't Generalize: The Limits of Detection in an Open World
Berisha, Visar
Kadambi, Prad
Lenz, Isabella
Cryptography and Security
Sound
Audio and Speech Processing
Speech deepfake detectors are often evaluated on clean, benchmark-style conditions, but deployment occurs in an open world of shifting devices, sampling rates, codecs, environments, and attack families. This creates a ``coverage debt" for AI-based detectors: every new condition multiplies with existing ones, producing data blind spots that grow faster than data can be collected. Because attackers can target these uncovered regions, worst-case performance (not average benchmark scores) determines security. To demonstrate the impact of the coverage debt problem, we analyze results from a recent cross-testing framework. Grouping performance by bona fide domain and spoof release year, two patterns emerge: newer synthesizers erase the legacy artifacts detectors rely on, and conversational speech domains (teleconferencing, interviews, social media) are consistently the hardest to secure. These findings show that detection alone should not be relied upon for high-stakes decisions. Detectors should be treated as auxiliary signals within layered defenses that include provenance, personhood credentials, and policy safeguards.
title Why Speech Deepfake Detectors Won't Generalize: The Limits of Detection in an Open World
topic Cryptography and Security
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.20405