ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mahapatra, Aurosweta, Ulgen, Ismail Rasim, Lee, Kong Aik, Andrews, Nicholas, Sisman, Berrak
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917409420476416
author Mahapatra, Aurosweta
Ulgen, Ismail Rasim
Lee, Kong Aik
Andrews, Nicholas
Sisman, Berrak
author_facet Mahapatra, Aurosweta
Ulgen, Ismail Rasim
Lee, Kong Aik
Andrews, Nicholas
Sisman, Berrak
contents Speech deepfake detection (SDD) systems perform well on standard benchmarks datasets but often fail to generalize to expressive and emotional spoofing attacks. Many methods rely on spoof-heavy training data, learning dataset-specific artifacts rather than transferable cues of natural speech. In contrast, humans internalize variability in real speech and detect fakes as deviations from it. We introduce ProSDD, a two-stage framework that enriches model embeddings through supervised masked prediction of speaker-conditioned prosodic variation based on pitch, voice activity, and energy. Stage I learns prosodic variability from real speech, and Stage II jointly optimizes this objective with spoof classification. ProSDD consistently outperforms baselines under both ASVspoof 2019 and 2024 training, reducing ASVspoof 2024 EER from 25.43% to 16.14% (2019-trained) and from 39.62% to 7.38% (2024-trained), while achieving 50% relative reductions on EmoFake and EmoSpoof-TTS.
format Preprint
id arxiv_https___arxiv_org_abs_2604_13229
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks
Mahapatra, Aurosweta
Ulgen, Ismail Rasim
Lee, Kong Aik
Andrews, Nicholas
Sisman, Berrak
Audio and Speech Processing
Speech deepfake detection (SDD) systems perform well on standard benchmarks datasets but often fail to generalize to expressive and emotional spoofing attacks. Many methods rely on spoof-heavy training data, learning dataset-specific artifacts rather than transferable cues of natural speech. In contrast, humans internalize variability in real speech and detect fakes as deviations from it. We introduce ProSDD, a two-stage framework that enriches model embeddings through supervised masked prediction of speaker-conditioned prosodic variation based on pitch, voice activity, and energy. Stage I learns prosodic variability from real speech, and Stage II jointly optimizes this objective with spoof classification. ProSDD consistently outperforms baselines under both ASVspoof 2019 and 2024 training, reducing ASVspoof 2024 EER from 25.43% to 16.14% (2019-trained) and from 39.62% to 7.38% (2024-trained), while achieving 50% relative reductions on EmoFake and EmoSpoof-TTS.
title ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks
topic Audio and Speech Processing
url https://arxiv.org/abs/2604.13229