X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Rixi, Liu, Qingyu, Li, Haitao, Chen, Yushen, Niu, Zhikang, Yang, Yunting, Zhao, Jian, Li, Ke, Sisman, Berrak, Cheng, Qinyuan, Qiu, Xipeng, Yu, Kai, Chen, Xie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
von: Li, Xuyuan, et al.
Veröffentlicht: (2024)
von: Li, Xuyuan, et al.
Veröffentlicht: (2024)
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
von: Du, Zongyang, et al.
Veröffentlicht: (2024)
von: Du, Zongyang, et al.
Veröffentlicht: (2024)
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
von: Lee, Philip H., et al.
Veröffentlicht: (2024)
von: Lee, Philip H., et al.
Veröffentlicht: (2024)
X-VC: Zero-shot Streaming Voice Conversion in Codec Space
von: Zheng, Qixi, et al.
Veröffentlicht: (2026)
von: Zheng, Qixi, et al.
Veröffentlicht: (2026)
Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis
von: Liu, Qingyu, et al.
Veröffentlicht: (2025)
von: Liu, Qingyu, et al.
Veröffentlicht: (2025)
Towards Naturalistic Voice Conversion: NaturalVoices Dataset with an Automatic Processing Pipeline
von: Salman, Ali N., et al.
Veröffentlicht: (2024)
von: Salman, Ali N., et al.
Veröffentlicht: (2024)
One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech
von: Abebe, Amanuel Gizachew, et al.
Veröffentlicht: (2026)
von: Abebe, Amanuel Gizachew, et al.
Veröffentlicht: (2026)
CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning
von: Li, Renyuan, et al.
Veröffentlicht: (2025)
von: Li, Renyuan, et al.
Veröffentlicht: (2025)
PRESENT: Zero-Shot Text-to-Prosody Control
von: Lam, Perry, et al.
Veröffentlicht: (2024)
von: Lam, Perry, et al.
Veröffentlicht: (2024)
Multi-modal Adversarial Training for Zero-Shot Voice Cloning
von: Janiczek, John, et al.
Veröffentlicht: (2024)
von: Janiczek, John, et al.
Veröffentlicht: (2024)
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2026)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2026)
Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
von: Dai, Shuqi, et al.
Veröffentlicht: (2025)
von: Dai, Shuqi, et al.
Veröffentlicht: (2025)
VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents
von: Li, Haiyun, et al.
Veröffentlicht: (2025)
von: Li, Haiyun, et al.
Veröffentlicht: (2025)
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
von: Du, Zongyang, et al.
Veröffentlicht: (2025)
von: Du, Zongyang, et al.
Veröffentlicht: (2025)
The THU-HCSI Multi-Speaker Multi-Lingual Few-Shot Voice Cloning System for LIMMITS'24 Challenge
von: Zhou, Yixuan, et al.
Veröffentlicht: (2024)
von: Zhou, Yixuan, et al.
Veröffentlicht: (2024)
SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
von: Lam, Perry, et al.
Veröffentlicht: (2022)
von: Lam, Perry, et al.
Veröffentlicht: (2022)
OpenVoice: Versatile Instant Voice Cloning
von: Qin, Zengyi, et al.
Veröffentlicht: (2023)
von: Qin, Zengyi, et al.
Veröffentlicht: (2023)
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
RefXVC: Cross-Lingual Voice Conversion with Enhanced Reference Leveraging
von: Zhang, Mingyang, et al.
Veröffentlicht: (2024)
von: Zhang, Mingyang, et al.
Veröffentlicht: (2024)
Who is Speaking or Who is Depressed? A Controlled Study of Speaker Leakage in Speech-Based Depression Detection
von: Yeh, Hsiang-Chen, et al.
Veröffentlicht: (2026)
von: Yeh, Hsiang-Chen, et al.
Veröffentlicht: (2026)
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue
von: Li, Ruiqi, et al.
Veröffentlicht: (2026)
von: Li, Ruiqi, et al.
Veröffentlicht: (2026)
Voice Cloning: Comprehensive Survey
von: Azzuni, Hussam, et al.
Veröffentlicht: (2025)
von: Azzuni, Hussam, et al.
Veröffentlicht: (2025)
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
von: Guan, Wenhao, et al.
Veröffentlicht: (2025)
von: Guan, Wenhao, et al.
Veröffentlicht: (2025)
MM-Sonate: Multimodal Controllable Audio-Video Generation with Zero-Shot Voice Cloning
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
von: Li, Junjie, et al.
Veröffentlicht: (2023)
von: Li, Junjie, et al.
Veröffentlicht: (2023)
ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed
von: Chen, Meiying, et al.
Veröffentlicht: (2022)
von: Chen, Meiying, et al.
Veröffentlicht: (2022)
GenVC: Self-Supervised Zero-Shot Voice Conversion
von: Cai, Zexin, et al.
Veröffentlicht: (2025)
von: Cai, Zexin, et al.
Veröffentlicht: (2025)
VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching
von: Choi, Ha-Yeong, et al.
Veröffentlicht: (2025)
von: Choi, Ha-Yeong, et al.
Veröffentlicht: (2025)
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
Zero-Shot Duet Singing Voices Separation with Diffusion Models
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2023)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2023)
Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
von: Niu, Zhikang, et al.
Veröffentlicht: (2025)
von: Niu, Zhikang, et al.
Veröffentlicht: (2025)
Enhancing Polyglot Voices by Leveraging Cross-Lingual Fine-Tuning in Any-to-One Voice Conversion
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2024)
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2024)
Exploring speech style spaces with language models: Emotional TTS without emotion labels
von: Chandra, Shreeram Suresh, et al.
Veröffentlicht: (2024)
von: Chandra, Shreeram Suresh, et al.
Veröffentlicht: (2024)
VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions
von: Zhan, Jun, et al.
Veröffentlicht: (2025)
von: Zhan, Jun, et al.
Veröffentlicht: (2025)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
von: Li, Xuyuan, et al.
Veröffentlicht: (2024) -
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
von: Du, Zongyang, et al.
Veröffentlicht: (2024) -
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
von: Lee, Philip H., et al.
Veröffentlicht: (2024) -
X-VC: Zero-shot Streaming Voice Conversion in Codec Space
von: Zheng, Qixi, et al.
Veröffentlicht: (2026) -
Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis
von: Liu, Qingyu, et al.
Veröffentlicht: (2025)