Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech Edits

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Sung-Feng, Kuo, Heng-Cheng, Chen, Zhehuai, Yang, Xuesong, Yang, Chao-Han Huck, Tsao, Yu, Wang, Yu-Chiang Frank, Lee, Hung-yi, Fu, Szu-Wei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929663690932224
author Huang, Sung-Feng
Kuo, Heng-Cheng
Chen, Zhehuai
Yang, Xuesong
Yang, Chao-Han Huck
Tsao, Yu
Wang, Yu-Chiang Frank
Lee, Hung-yi
Fu, Szu-Wei
author_facet Huang, Sung-Feng
Kuo, Heng-Cheng
Chen, Zhehuai
Yang, Xuesong
Yang, Chao-Han Huck
Tsao, Yu
Wang, Yu-Chiang Frank
Lee, Hung-yi
Fu, Szu-Wei
contents Neural speech editing advancements have raised concerns about their misuse in spoofing attacks. Traditional partially edited speech corpora primarily focus on cut-and-paste edits, which, while maintaining speaker consistency, often introduce detectable discontinuities. Recent methods, like A\textsuperscript{3}T and Voicebox, improve transitions by leveraging contextual information. To foster spoofing detection research, we introduce the Speech INfilling Edit (SINE) dataset, created with Voicebox. We detailed the process of re-implementing Voicebox training and dataset creation. Subjective evaluations confirm that speech edited using this novel technique is more challenging to detect than conventional cut-and-paste methods. Despite human difficulty, experimental results demonstrate that self-supervised-based detectors can achieve remarkable performance in detection, localization, and generalization across different edit methods. The dataset and related models will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2501_03805
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech Edits
Huang, Sung-Feng
Kuo, Heng-Cheng
Chen, Zhehuai
Yang, Xuesong
Yang, Chao-Han Huck
Tsao, Yu
Wang, Yu-Chiang Frank
Lee, Hung-yi
Fu, Szu-Wei
Sound
Computation and Language
Audio and Speech Processing
Neural speech editing advancements have raised concerns about their misuse in spoofing attacks. Traditional partially edited speech corpora primarily focus on cut-and-paste edits, which, while maintaining speaker consistency, often introduce detectable discontinuities. Recent methods, like A\textsuperscript{3}T and Voicebox, improve transitions by leveraging contextual information. To foster spoofing detection research, we introduce the Speech INfilling Edit (SINE) dataset, created with Voicebox. We detailed the process of re-implementing Voicebox training and dataset creation. Subjective evaluations confirm that speech edited using this novel technique is more challenging to detect than conventional cut-and-paste methods. Despite human difficulty, experimental results demonstrate that self-supervised-based detectors can achieve remarkable performance in detection, localization, and generalization across different edit methods. The dataset and related models will be made publicly available.
title Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech Edits
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2501.03805