Improving Generalization for AI-Synthesized Voice Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ren, Hainan, Lin, Li, Liu, Chun-Hao, Wang, Xin, Hu, Shu
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910766099070976
author Ren, Hainan
Lin, Li
Liu, Chun-Hao
Wang, Xin
Hu, Shu
author_facet Ren, Hainan
Lin, Li
Liu, Chun-Hao
Wang, Xin
Hu, Shu
contents AI-synthesized voice technology has the potential to create realistic human voices for beneficial applications, but it can also be misused for malicious purposes. While existing AI-synthesized voice detection models excel in intra-domain evaluation, they face challenges in generalizing across different domains, potentially becoming obsolete as new voice generators emerge. Current solutions use diverse data and advanced machine learning techniques (e.g., domain-invariant representation, self-supervised learning), but are limited by predefined vocoders and sensitivity to factors like background noise and speaker identity. In this work, we introduce an innovative disentanglement framework aimed at extracting domain-agnostic artifact features related to vocoders. Utilizing these features, we enhance model learning in a flat loss landscape, enabling escape from suboptimal solutions and improving generalization. Extensive experiments on benchmarks show our approach outperforms state-of-the-art methods, achieving up to 5.12% improvement in the equal error rate metric in intra-domain and 7.59% in cross-domain evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19279
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Generalization for AI-Synthesized Voice Detection
Ren, Hainan
Lin, Li
Liu, Chun-Hao
Wang, Xin
Hu, Shu
Sound
Machine Learning
Audio and Speech Processing
AI-synthesized voice technology has the potential to create realistic human voices for beneficial applications, but it can also be misused for malicious purposes. While existing AI-synthesized voice detection models excel in intra-domain evaluation, they face challenges in generalizing across different domains, potentially becoming obsolete as new voice generators emerge. Current solutions use diverse data and advanced machine learning techniques (e.g., domain-invariant representation, self-supervised learning), but are limited by predefined vocoders and sensitivity to factors like background noise and speaker identity. In this work, we introduce an innovative disentanglement framework aimed at extracting domain-agnostic artifact features related to vocoders. Utilizing these features, we enhance model learning in a flat loss landscape, enabling escape from suboptimal solutions and improving generalization. Extensive experiments on benchmarks show our approach outperforms state-of-the-art methods, achieving up to 5.12% improvement in the equal error rate metric in intra-domain and 7.59% in cross-domain evaluations.
title Improving Generalization for AI-Synthesized Voice Detection
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2412.19279