Voice Cloning: Comprehensive Survey
Fuente:
arXiv
Saved in:
| Main Authors: | Azzuni, Hussam, Saddik, Abdulmotaleb El |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning
by: Li, Renyuan, et al.
Published: (2025)
by: Li, Renyuan, et al.
Published: (2025)
Pronunciation Deviation Analysis Through Voice Cloning and Acoustic Comparison
by: Valdivia, Andrew, et al.
Published: (2025)
by: Valdivia, Andrew, et al.
Published: (2025)
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
by: Xu, Rixi, et al.
Published: (2026)
by: Xu, Rixi, et al.
Published: (2026)
Fed-PISA: Federated Voice Cloning via Personalized Identity-Style Adaptation
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
by: Moell, Birger, et al.
Published: (2025)
by: Moell, Birger, et al.
Published: (2025)
VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents
by: Li, Haiyun, et al.
Published: (2025)
by: Li, Haiyun, et al.
Published: (2025)
VoiceGRPO: Modern MoE Transformers with Group Relative Policy Optimization GRPO for AI Voice Health Care Applications on Voice Pathology Detection
by: Togootogtokh, Enkhtogtokh, et al.
Published: (2025)
by: Togootogtokh, Enkhtogtokh, et al.
Published: (2025)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
by: Anastassiou, Philip, et al.
Published: (2024)
by: Anastassiou, Philip, et al.
Published: (2024)
Voice Attribute Editing with Text Prompt
by: Sheng, Zhengyan, et al.
Published: (2024)
by: Sheng, Zhengyan, et al.
Published: (2024)
A New Approach to Voice Authenticity
by: Müller, Nicolas M., et al.
Published: (2024)
by: Müller, Nicolas M., et al.
Published: (2024)
Deepfake Detection of Singing Voices With Whisper Encodings
by: Sharma, Falguni, et al.
Published: (2025)
by: Sharma, Falguni, et al.
Published: (2025)
SingFake: Singing Voice Deepfake Detection
by: Zang, Yongyi, et al.
Published: (2023)
by: Zang, Yongyi, et al.
Published: (2023)
Degrading Voice: A Comprehensive Overview of Robust Voice Conversion Through Input Manipulation
by: Song, Xining, et al.
Published: (2025)
by: Song, Xining, et al.
Published: (2025)
LTS-VoiceAgent: A Listen-Think-Speak Framework for Efficient Streaming Voice Interaction via Semantic Triggering and Incremental Reasoning
by: Zou, Wenhao, et al.
Published: (2026)
by: Zou, Wenhao, et al.
Published: (2026)
Physics-Guided Deepfake Detection for Voice Authentication Systems
by: Mohammadi, Alireza, et al.
Published: (2025)
by: Mohammadi, Alireza, et al.
Published: (2025)
The Voice Timbre Attribute Detection 2025 Challenge Evaluation Plan
by: Sheng, Zhengyan, et al.
Published: (2025)
by: Sheng, Zhengyan, et al.
Published: (2025)
MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
by: Huang, Jiawei, et al.
Published: (2024)
by: Huang, Jiawei, et al.
Published: (2024)
Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
by: Wang, Rui, et al.
Published: (2024)
by: Wang, Rui, et al.
Published: (2024)
RDSinger: Reference-based Diffusion Network for Singing Voice Synthesis
by: Sui, Kehan, et al.
Published: (2024)
by: Sui, Kehan, et al.
Published: (2024)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
by: Li, Xuyuan, et al.
Published: (2024)
by: Li, Xuyuan, et al.
Published: (2024)
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
by: Zhang, Xueyao, et al.
Published: (2025)
by: Zhang, Xueyao, et al.
Published: (2025)
Controllable Singing Voice Synthesis using Phoneme-Level Energy Sequence
by: Ryu, Yerin, et al.
Published: (2025)
by: Ryu, Yerin, et al.
Published: (2025)
EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion
by: Joglekar, Advait, et al.
Published: (2025)
by: Joglekar, Advait, et al.
Published: (2025)
Quantum-Inspired Audio Unlearning: Towards Privacy-Preserving Voice Biometrics
by: Pathak, Shreyansh, et al.
Published: (2025)
by: Pathak, Shreyansh, et al.
Published: (2025)
SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
by: Bai, Bingsong, et al.
Published: (2024)
by: Bai, Bingsong, et al.
Published: (2024)
Improving Voice Quality in Speech Anonymization With Just Perception-Informed Losses
by: Ghosh, Suhita, et al.
Published: (2024)
by: Ghosh, Suhita, et al.
Published: (2024)
Spectral Mapping of Singing Voices: U-Net-Assisted Vocal Segmentation
by: Sorrenti, Adam
Published: (2024)
by: Sorrenti, Adam
Published: (2024)
LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling
by: Huang, Yubo, et al.
Published: (2024)
by: Huang, Yubo, et al.
Published: (2024)
A Real-Time Voice Activity Detection Based On Lightweight Neural
by: Jia, Jidong, et al.
Published: (2024)
by: Jia, Jidong, et al.
Published: (2024)
A Novel Labeled Human Voice Signal Dataset for Misbehavior Detection
by: Raza, Ali, et al.
Published: (2024)
by: Raza, Ali, et al.
Published: (2024)
Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders
by: Lau, Hok-Shing, et al.
Published: (2024)
by: Lau, Hok-Shing, et al.
Published: (2024)
SelfVC: Voice Conversion With Iterative Refinement using Self Transformations
by: Neekhara, Paarth, et al.
Published: (2023)
by: Neekhara, Paarth, et al.
Published: (2023)
Contrastive Learning-based Chaining-Cluster for Multilingual Voice-Face Association
by: Chen, Wuyang, et al.
Published: (2024)
by: Chen, Wuyang, et al.
Published: (2024)
VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching
by: Choi, Ha-Yeong, et al.
Published: (2025)
by: Choi, Ha-Yeong, et al.
Published: (2025)
Generative Adversarial Network based Voice Conversion: Techniques, Challenges, and Recent Advancements
by: Dhar, Sandipan, et al.
Published: (2025)
by: Dhar, Sandipan, et al.
Published: (2025)
VoiceBridge: General Speech Restoration with One-step Latent Bridge Models
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
by: Akti, Seymanur, et al.
Published: (2025)
by: Akti, Seymanur, et al.
Published: (2025)
DiTSinger: Scaling Singing Voice Synthesis with Diffusion Transformer and Implicit Alignment
by: Du, Zongcai, et al.
Published: (2025)
by: Du, Zongcai, et al.
Published: (2025)
Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion
by: Chen, Yun, et al.
Published: (2023)
by: Chen, Yun, et al.
Published: (2023)
Fast-VGAN: Lightweight Voice Conversion with Explicit Control of F0 and Duration Parameters
by: Abrassart, Mathilde, et al.
Published: (2025)
by: Abrassart, Mathilde, et al.
Published: (2025)
Similar Items
-
CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning
by: Li, Renyuan, et al.
Published: (2025) -
Pronunciation Deviation Analysis Through Voice Cloning and Acoustic Comparison
by: Valdivia, Andrew, et al.
Published: (2025) -
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
by: Xu, Rixi, et al.
Published: (2026) -
Fed-PISA: Federated Voice Cloning via Personalized Identity-Style Adaptation
by: Wang, Qi, et al.
Published: (2025) -
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
by: Moell, Birger, et al.
Published: (2025)