Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chinchmalatpure, Prajwal, Chinchmalatpure, Suyash, Chavan, Siddharth
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:https://arxiv.org/abs/2601.04227
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909984810336256
author Chinchmalatpure, Prajwal
Chinchmalatpure, Suyash
Chavan, Siddharth
author_facet Chinchmalatpure, Prajwal
Chinchmalatpure, Suyash
Chavan, Siddharth
contents Generative audio technologies now enable highly realistic voice cloning and real-time voice conversion, increasing the risk of impersonation, fraud, and misinformation in communication channels such as phone and video calls. This study investigates real-time detection of AI-generated speech produced using Retrieval-based Voice Conversion (RVC), evaluated on the DEEP-VOICE dataset, which includes authentic and voice-converted speech samples from multiple well-known speakers. To simulate realistic conditions, deepfake generation is applied to isolated vocal components, followed by the reintroduction of background ambiance to suppress trivial artifacts and emphasize conversion-specific cues. We frame detection as a streaming classification task by dividing audio into one-second segments, extracting time-frequency and cepstral features, and training supervised machine learning models to classify each segment as real or voice-converted. The proposed system enables low-latency inference, supporting both segment-level decisions and call-level aggregation. Experimental results show that short-window acoustic features can reliably capture discriminative patterns associated with RVC speech, even in noisy backgrounds. These findings demonstrate the feasibility of practical, real-time deepfake speech detection and underscore the importance of evaluating under realistic audio mixing conditions for robust deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2601_04227
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Defense Against Synthetic Speech: Real-Time Detection of RVC Voice Conversion Attacks
Chinchmalatpure, Prajwal
Chinchmalatpure, Suyash
Chavan, Siddharth
Sound
Artificial Intelligence
Audio and Speech Processing
Generative audio technologies now enable highly realistic voice cloning and real-time voice conversion, increasing the risk of impersonation, fraud, and misinformation in communication channels such as phone and video calls. This study investigates real-time detection of AI-generated speech produced using Retrieval-based Voice Conversion (RVC), evaluated on the DEEP-VOICE dataset, which includes authentic and voice-converted speech samples from multiple well-known speakers. To simulate realistic conditions, deepfake generation is applied to isolated vocal components, followed by the reintroduction of background ambiance to suppress trivial artifacts and emphasize conversion-specific cues. We frame detection as a streaming classification task by dividing audio into one-second segments, extracting time-frequency and cepstral features, and training supervised machine learning models to classify each segment as real or voice-converted. The proposed system enables low-latency inference, supporting both segment-level decisions and call-level aggregation. Experimental results show that short-window acoustic features can reliably capture discriminative patterns associated with RVC speech, even in noisy backgrounds. These findings demonstrate the feasibility of practical, real-time deepfake speech detection and underscore the importance of evaluating under realistic audio mixing conditions for robust deployment.
title Defense Against Synthetic Speech: Real-Time Detection of RVC Voice Conversion Attacks
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2601.04227