Improving Generalization for AI-Synthesized Voice Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Ren, Hainan, Lin, Li, Liu, Chun-Hao, Wang, Xin, Hu, Shu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OpenVoice: Versatile Instant Voice Cloning
by: Qin, Zengyi, et al.
Published: (2023)
by: Qin, Zengyi, et al.
Published: (2023)
SEF-MK: Speaker-Embedding-Free Voice Anonymization through Multi-k-means Quantization
by: Tang, Beilong, et al.
Published: (2025)
by: Tang, Beilong, et al.
Published: (2025)
SiFiSinger: A High-Fidelity End-to-End Singing Voice Synthesizer based on Source-filter Model
by: Cui, Jianwei, et al.
Published: (2024)
by: Cui, Jianwei, et al.
Published: (2024)
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
by: Lee, Philip H., et al.
Published: (2024)
by: Lee, Philip H., et al.
Published: (2024)
SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models
by: Yin, Chun, et al.
Published: (2024)
by: Yin, Chun, et al.
Published: (2024)
BiSinger: Bilingual Singing Voice Synthesis
by: Zhou, Huali, et al.
Published: (2023)
by: Zhou, Huali, et al.
Published: (2023)
A Concept-based approach to Voice Disorder Detection
by: Ghia, Davide, et al.
Published: (2025)
by: Ghia, Davide, et al.
Published: (2025)
On the Generation and Removal of Speaker Adversarial Perturbation for Voice-Privacy Protection
by: Guo, Chenyang, et al.
Published: (2024)
by: Guo, Chenyang, et al.
Published: (2024)
Creative Text-to-Audio Generation via Synthesizer Programming
by: Cherep, Manuel, et al.
Published: (2024)
by: Cherep, Manuel, et al.
Published: (2024)
An AI-Driven Approach to Wind Turbine Bearing Fault Diagnosis from Acoustic Signals
by: Wang, Zhao, et al.
Published: (2024)
by: Wang, Zhao, et al.
Published: (2024)
Multi-blank Transducers for Speech Recognition
by: Xu, Hainan, et al.
Published: (2022)
by: Xu, Hainan, et al.
Published: (2022)
Zero-shot Voice Conversion with Diffusion Transformers
by: Liu, Songting
Published: (2024)
by: Liu, Songting
Published: (2024)
Multi-modal Adversarial Training for Zero-Shot Voice Cloning
by: Janiczek, John, et al.
Published: (2024)
by: Janiczek, John, et al.
Published: (2024)
ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps
by: Song, Yulin, et al.
Published: (2024)
by: Song, Yulin, et al.
Published: (2024)
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
by: Du, Zongyang, et al.
Published: (2025)
by: Du, Zongyang, et al.
Published: (2025)
Diffusion Synthesizer for Efficient Multilingual Speech to Speech Translation
by: Hirschkind, Nameer, et al.
Published: (2024)
by: Hirschkind, Nameer, et al.
Published: (2024)
StreamVC: Real-Time Low-Latency Voice Conversion
by: Yang, Yang, et al.
Published: (2024)
by: Yang, Yang, et al.
Published: (2024)
End-to-End Integration of Speech Separation and Voice Activity Detection for Low-Latency Diarization of Telephone Conversations
by: Morrone, Giovanni, et al.
Published: (2023)
by: Morrone, Giovanni, et al.
Published: (2023)
VANPY: Voice Analysis Framework
by: Koushnir, Gregory, et al.
Published: (2025)
by: Koushnir, Gregory, et al.
Published: (2025)
Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect
by: Narain, Jaya, et al.
Published: (2025)
by: Narain, Jaya, et al.
Published: (2025)
Improving Real-Time Music Accompaniment Separation with MMDenseNet
by: Wang, Chun-Hsiang, et al.
Published: (2024)
by: Wang, Chun-Hsiang, et al.
Published: (2024)
Towards General-Purpose Text-Instruction-Guided Voice Conversion
by: Kuan, Chun-Yi, et al.
Published: (2023)
by: Kuan, Chun-Yi, et al.
Published: (2023)
Utilizing TTS Synthesized Data for Efficient Development of Keyword Spotting Model
by: Park, Hyun Jin, et al.
Published: (2024)
by: Park, Hyun Jin, et al.
Published: (2024)
Compact Neural TTS Voices for Accessibility
by: Jain, Kunal, et al.
Published: (2025)
by: Jain, Kunal, et al.
Published: (2025)
Speech to Speech Synthesis for Voice Impersonation
by: Johnson, Bjorn, et al.
Published: (2026)
by: Johnson, Bjorn, et al.
Published: (2026)
Discrete Optimal Transport and Voice Conversion
by: Selitskiy, Anton, et al.
Published: (2025)
by: Selitskiy, Anton, et al.
Published: (2025)
Phoneme Hallucinator: One-shot Voice Conversion via Set Expansion
by: Shan, Siyuan, et al.
Published: (2023)
by: Shan, Siyuan, et al.
Published: (2023)
Synthesizer Sound Matching Using Audio Spectrogram Transformers
by: Bruford, Fred, et al.
Published: (2024)
by: Bruford, Fred, et al.
Published: (2024)
Optimal Transport Maps are Good Voice Converters
by: Asadulaev, Arip, et al.
Published: (2024)
by: Asadulaev, Arip, et al.
Published: (2024)
Tessellated Linear Model for Age Prediction from Voice
by: Alharthi, Dareen, et al.
Published: (2025)
by: Alharthi, Dareen, et al.
Published: (2025)
Singing Voice Conversion with Accompaniment Using Self-Supervised Representation-Based Melody Features
by: Chen, Wei, et al.
Published: (2025)
by: Chen, Wei, et al.
Published: (2025)
Mouth Articulation-Based Anchoring for Improved Cross-Corpus Speech Emotion Recognition
by: Upadhyay, Shreya G., et al.
Published: (2024)
by: Upadhyay, Shreya G., et al.
Published: (2024)
Unified AI for Accurate Audio Anomaly Detection
by: Khaleghpour, Hamideh, et al.
Published: (2025)
by: Khaleghpour, Hamideh, et al.
Published: (2025)
Voice Signal Processing for Machine Learning. The Case of Speaker Isolation
by: Ganchev, Radan
Published: (2024)
by: Ganchev, Radan
Published: (2024)
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
by: Ulgen, Ismail Rasim, et al.
Published: (2026)
by: Ulgen, Ismail Rasim, et al.
Published: (2026)
Towards Robust Assessment of Pathological Voices via Combined Low-Level Descriptors and Foundation Model Representations
by: Ariyanti, Whenty, et al.
Published: (2025)
by: Ariyanti, Whenty, et al.
Published: (2025)
Voice Conversion with Diverse Intonation using Conditional Variational Auto-Encoder
by: Suh, Soobin, et al.
Published: (2025)
by: Suh, Soobin, et al.
Published: (2025)
An Enhanced Audio Feature Tailored for Anomalous Sound Detection Based on Pre-trained Models
by: Zhong, Guirui, et al.
Published: (2025)
by: Zhong, Guirui, et al.
Published: (2025)
Evaluating Echo State Network for Parkinson's Disease Prediction using Voice Features
by: Hosseininian, Seyedeh Zahra Seyedi, et al.
Published: (2024)
by: Hosseininian, Seyedeh Zahra Seyedi, et al.
Published: (2024)
Does Audio Deepfake Detection Generalize?
by: Müller, Nicolas M., et al.
Published: (2022)
by: Müller, Nicolas M., et al.
Published: (2022)
Similar Items
-
OpenVoice: Versatile Instant Voice Cloning
by: Qin, Zengyi, et al.
Published: (2023) -
SEF-MK: Speaker-Embedding-Free Voice Anonymization through Multi-k-means Quantization
by: Tang, Beilong, et al.
Published: (2025) -
SiFiSinger: A High-Fidelity End-to-End Singing Voice Synthesizer based on Source-filter Model
by: Cui, Jianwei, et al.
Published: (2024) -
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
by: Lee, Philip H., et al.
Published: (2024) -
SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models
by: Yin, Chun, et al.
Published: (2024)