Alethia: A Foundational Encoder for Voice Deepfakes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Yi, Dwivedi, Brahmi, Raghuram, Jayaram, Koppisetti, Surya
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911637951217664
author Zhu, Yi
Dwivedi, Brahmi
Raghuram, Jayaram
Koppisetti, Surya
author_facet Zhu, Yi
Dwivedi, Brahmi
Raghuram, Jayaram
Koppisetti, Surya
contents Existing voice deepfake detection and localization models rely heavily on representations extracted from speech foundation models (SFMs). However, downstream finetuning has now reached a state of diminishing returns. In this paper, we shift the focus to pretraining and propose a novel recipe that combines bottleneck masked embedding prediction with flow-matching based spectrogram reconstruction. The outcome, Alethia, is the first foundational audio encoder for various voice deepfake detection and localization tasks. We evaluate on $5$ different tasks with $56$ benchmark datasets, and note Alethia significantly outperforms state-of-the-art SFMs with superior robustness to real-world perturbations and zero-shot generalization to unseen domains (e.g., singing deepfakes). We also demonstrate the limitation of discrete targets in masked token prediction, and show the importance of continuous embedding prediction and generative pretraining for capturing deepfake artifacts.
format Preprint
id arxiv_https___arxiv_org_abs_2605_00251
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Alethia: A Foundational Encoder for Voice Deepfakes
Zhu, Yi
Dwivedi, Brahmi
Raghuram, Jayaram
Koppisetti, Surya
Sound
Computation and Language
Audio and Speech Processing
Existing voice deepfake detection and localization models rely heavily on representations extracted from speech foundation models (SFMs). However, downstream finetuning has now reached a state of diminishing returns. In this paper, we shift the focus to pretraining and propose a novel recipe that combines bottleneck masked embedding prediction with flow-matching based spectrogram reconstruction. The outcome, Alethia, is the first foundational audio encoder for various voice deepfake detection and localization tasks. We evaluate on $5$ different tasks with $56$ benchmark datasets, and note Alethia significantly outperforms state-of-the-art SFMs with superior robustness to real-world perturbations and zero-shot generalization to unseen domains (e.g., singing deepfakes). We also demonstrate the limitation of discrete targets in masked token prediction, and show the importance of continuous embedding prediction and generative pretraining for capturing deepfake artifacts.
title Alethia: A Foundational Encoder for Voice Deepfakes
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2605.00251