Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech Edits
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866929663690932224 |
|---|---|
| author | Huang, Sung-Feng Kuo, Heng-Cheng Chen, Zhehuai Yang, Xuesong Yang, Chao-Han Huck Tsao, Yu Wang, Yu-Chiang Frank Lee, Hung-yi Fu, Szu-Wei |
| author_facet | Huang, Sung-Feng Kuo, Heng-Cheng Chen, Zhehuai Yang, Xuesong Yang, Chao-Han Huck Tsao, Yu Wang, Yu-Chiang Frank Lee, Hung-yi Fu, Szu-Wei |
| contents | Neural speech editing advancements have raised concerns about their misuse in spoofing attacks. Traditional partially edited speech corpora primarily focus on cut-and-paste edits, which, while maintaining speaker consistency, often introduce detectable discontinuities. Recent methods, like A\textsuperscript{3}T and Voicebox, improve transitions by leveraging contextual information. To foster spoofing detection research, we introduce the Speech INfilling Edit (SINE) dataset, created with Voicebox. We detailed the process of re-implementing Voicebox training and dataset creation. Subjective evaluations confirm that speech edited using this novel technique is more challenging to detect than conventional cut-and-paste methods. Despite human difficulty, experimental results demonstrate that self-supervised-based detectors can achieve remarkable performance in detection, localization, and generalization across different edit methods. The dataset and related models will be made publicly available. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_03805 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech Edits Huang, Sung-Feng Kuo, Heng-Cheng Chen, Zhehuai Yang, Xuesong Yang, Chao-Han Huck Tsao, Yu Wang, Yu-Chiang Frank Lee, Hung-yi Fu, Szu-Wei Sound Computation and Language Audio and Speech Processing Neural speech editing advancements have raised concerns about their misuse in spoofing attacks. Traditional partially edited speech corpora primarily focus on cut-and-paste edits, which, while maintaining speaker consistency, often introduce detectable discontinuities. Recent methods, like A\textsuperscript{3}T and Voicebox, improve transitions by leveraging contextual information. To foster spoofing detection research, we introduce the Speech INfilling Edit (SINE) dataset, created with Voicebox. We detailed the process of re-implementing Voicebox training and dataset creation. Subjective evaluations confirm that speech edited using this novel technique is more challenging to detect than conventional cut-and-paste methods. Despite human difficulty, experimental results demonstrate that self-supervised-based detectors can achieve remarkable performance in detection, localization, and generalization across different edit methods. The dataset and related models will be made publicly available. |
| title | Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech Edits |
| topic | Sound Computation and Language Audio and Speech Processing |
| url | https://arxiv.org/abs/2501.03805 |