Deepfake Detection of Singing Voices With Whisper Encodings

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sharma, Falguni, Gupta, Priyanka
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910807159209984
author Sharma, Falguni
Gupta, Priyanka
author_facet Sharma, Falguni
Gupta, Priyanka
contents The deepfake generation of singing vocals is a concerning issue for artists in the music industry. In this work, we propose a singing voice deepfake detection (SVDD) system, which uses noise-variant encodings of open-AI's Whisper model. As counter-intuitive as it may sound, even though the Whisper model is known to be noise-robust, the encodings are rich in non-speech information, and are noise-variant. This leads us to evaluate Whisper encodings as feature representations for the SVDD task. Therefore, in this work, the SVDD task is performed on vocals and mixtures, and the performance is evaluated in \%EER over varying Whisper model sizes and two classifiers- CNN and ResNet34, under different testing conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18919
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Deepfake Detection of Singing Voices With Whisper Encodings
Sharma, Falguni
Gupta, Priyanka
Sound
Artificial Intelligence
Audio and Speech Processing
The deepfake generation of singing vocals is a concerning issue for artists in the music industry. In this work, we propose a singing voice deepfake detection (SVDD) system, which uses noise-variant encodings of open-AI's Whisper model. As counter-intuitive as it may sound, even though the Whisper model is known to be noise-robust, the encodings are rich in non-speech information, and are noise-variant. This leads us to evaluate Whisper encodings as feature representations for the SVDD task. Therefore, in this work, the SVDD task is performed on vocals and mixtures, and the performance is evaluated in \%EER over varying Whisper model sizes and two classifiers- CNN and ResNet34, under different testing conditions.
title Deepfake Detection of Singing Voices With Whisper Encodings
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2501.18919