Convert and Speak: Zero-shot Accent Conversion with Minimum Supervision
Fuente:
arXiv
Saved in:
| Main Authors: | Jia, Zhijun, Xue, Huaying, Peng, Xiulian, Lu, Yan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Low-latency Speech Enhancement via Speech Token Generation
by: Xue, Huaying, et al.
Published: (2023)
by: Xue, Huaying, et al.
Published: (2023)
Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
by: Jiang, Xue, et al.
Published: (2025)
by: Jiang, Xue, et al.
Published: (2025)
Latent-Domain Predictive Neural Speech Coding
by: Jiang, Xue, et al.
Published: (2022)
by: Jiang, Xue, et al.
Published: (2022)
Non-autoregressive real-time Accent Conversion model with voice cloning
by: Nechaev, Vladimir, et al.
Published: (2024)
by: Nechaev, Vladimir, et al.
Published: (2024)
EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion
by: Joglekar, Advait, et al.
Published: (2025)
by: Joglekar, Advait, et al.
Published: (2025)
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
by: Du, Zhihao, et al.
Published: (2024)
by: Du, Zhihao, et al.
Published: (2024)
Text-Queried Audio Source Separation via Hierarchical Modeling
by: Yin, Xinlei, et al.
Published: (2025)
by: Yin, Xinlei, et al.
Published: (2025)
CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data
by: Bai, Qibing, et al.
Published: (2026)
by: Bai, Qibing, et al.
Published: (2026)
Improving Pronunciation and Accent Conversion through Knowledge Distillation And Synthetic Ground-Truth from Native TTS
by: Nguyen, Tuan Nam, et al.
Published: (2024)
by: Nguyen, Tuan Nam, et al.
Published: (2024)
R2-SVC: Towards Real-World Robust and Expressive Zero-shot Singing Voice Conversion
by: Zheng, Junjie, et al.
Published: (2025)
by: Zheng, Junjie, et al.
Published: (2025)
Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
by: Kim, Jaeyoung, et al.
Published: (2024)
by: Kim, Jaeyoung, et al.
Published: (2024)
Towards Lightweight and Stable Zero-shot TTS with Self-distilled Representation Disentanglement
by: Chen, Qianniu, et al.
Published: (2025)
by: Chen, Qianniu, et al.
Published: (2025)
Controllable Accent Normalization via Discrete Diffusion
by: Bai, Qibing, et al.
Published: (2026)
by: Bai, Qibing, et al.
Published: (2026)
Masked Audio Modeling with CLAP and Multi-Objective Learning
by: Xin, Yifei, et al.
Published: (2024)
by: Xin, Yifei, et al.
Published: (2024)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
by: Melechovsky, Jan, et al.
Published: (2024)
by: Melechovsky, Jan, et al.
Published: (2024)
Selective Classifier-free Guidance for Zero-shot Text-to-speech
by: Zheng, John, et al.
Published: (2025)
by: Zheng, John, et al.
Published: (2025)
Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models
by: Chen, Sijing, et al.
Published: (2024)
by: Chen, Sijing, et al.
Published: (2024)
Zero-shot Musical Stem Retrieval with Joint-Embedding Predictive Architectures
by: Riou, Alain, et al.
Published: (2024)
by: Riou, Alain, et al.
Published: (2024)
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
by: Zhang, Xueyao, et al.
Published: (2025)
by: Zhang, Xueyao, et al.
Published: (2025)
ACES: Accent Subspaces for Coupling, Explanations, and Stress-Testing in Automatic Speech Recognition
by: Parekh, Swapnil
Published: (2026)
by: Parekh, Swapnil
Published: (2026)
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
by: Xu, Rixi, et al.
Published: (2026)
by: Xu, Rixi, et al.
Published: (2026)
DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning
by: Li, Xiquan, et al.
Published: (2024)
by: Li, Xiquan, et al.
Published: (2024)
OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching
by: Huynh-Nguyen, Hieu-Nghia, et al.
Published: (2025)
by: Huynh-Nguyen, Hieu-Nghia, et al.
Published: (2025)
Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition
by: Chen, Jinming, et al.
Published: (2024)
by: Chen, Jinming, et al.
Published: (2024)
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
by: Akti, Seymanur, et al.
Published: (2025)
by: Akti, Seymanur, et al.
Published: (2025)
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
by: Inoue, Sho, et al.
Published: (2024)
by: Inoue, Sho, et al.
Published: (2024)
GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech
by: Wang, Wenbin, et al.
Published: (2024)
by: Wang, Wenbin, et al.
Published: (2024)
QR-VC: Leveraging Quantization Residuals for Linear Disentanglement in Zero-Shot Voice Conversion
by: Sim, Youngjun, et al.
Published: (2024)
by: Sim, Youngjun, et al.
Published: (2024)
Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
by: Zhao, Junchuan, et al.
Published: (2025)
by: Zhao, Junchuan, et al.
Published: (2025)
Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching
by: Pan, Yu, et al.
Published: (2024)
by: Pan, Yu, et al.
Published: (2024)
Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
by: Wang, Kaidi, et al.
Published: (2025)
by: Wang, Kaidi, et al.
Published: (2025)
HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios
by: Bai, Bingsong, et al.
Published: (2025)
by: Bai, Bingsong, et al.
Published: (2025)
Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion
by: Wang, Zhichao, et al.
Published: (2023)
by: Wang, Zhichao, et al.
Published: (2023)
Serial-Parallel Dual-Path Architecture for Speaking Style Recognition
by: Li, Guojian, et al.
Published: (2025)
by: Li, Guojian, et al.
Published: (2025)
CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
by: Luong, Justin, et al.
Published: (2025)
by: Luong, Justin, et al.
Published: (2025)
Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling
by: Yang, Yuguang, et al.
Published: (2024)
by: Yang, Yuguang, et al.
Published: (2024)
DSFlow: Dual Supervision and Step-Aware Architecture for One-Step Flow Matching Speech Synthesis
by: Lin, Bin, et al.
Published: (2026)
by: Lin, Bin, et al.
Published: (2026)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
by: Anastassiou, Philip, et al.
Published: (2024)
by: Anastassiou, Philip, et al.
Published: (2024)
On the Relationship between Accent Strength and Articulatory Features
by: Huang, Kevin, et al.
Published: (2025)
by: Huang, Kevin, et al.
Published: (2025)
An Effective Automated Speaking Assessment Approach to Mitigating Data Scarcity and Imbalanced Distribution
by: Lo, Tien-Hong, et al.
Published: (2024)
by: Lo, Tien-Hong, et al.
Published: (2024)
Similar Items
-
Low-latency Speech Enhancement via Speech Token Generation
by: Xue, Huaying, et al.
Published: (2023) -
Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
by: Jiang, Xue, et al.
Published: (2025) -
Latent-Domain Predictive Neural Speech Coding
by: Jiang, Xue, et al.
Published: (2022) -
Non-autoregressive real-time Accent Conversion model with voice cloning
by: Nechaev, Vladimir, et al.
Published: (2024) -
EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion
by: Joglekar, Advait, et al.
Published: (2025)