Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yu, Tian, Baotong, Duan, Zhiyao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed
von: Chen, Meiying, et al.
Veröffentlicht: (2022)
von: Chen, Meiying, et al.
Veröffentlicht: (2022)
SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
von: Li, Junjie, et al.
Veröffentlicht: (2023)
von: Li, Junjie, et al.
Veröffentlicht: (2023)
PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
von: Zhang, You, et al.
Veröffentlicht: (2025)
von: Zhang, You, et al.
Veröffentlicht: (2025)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
Generating Novel and Realistic Speakers for Voice Conversion
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
AdaptVC: High Quality Voice Conversion with Adaptive Learning
von: Kim, Jaehun, et al.
Veröffentlicht: (2025)
von: Kim, Jaehun, et al.
Veröffentlicht: (2025)
Voice Impression Control in Zero-Shot TTS
von: Fujita, Kenichi, et al.
Veröffentlicht: (2025)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2025)
Not that Groove: Zero-Shot Symbolic Music Editing
von: Zhang, Li
Veröffentlicht: (2025)
von: Zhang, Li
Veröffentlicht: (2025)
TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions
von: Kang, Wonjune, et al.
Veröffentlicht: (2022)
von: Kang, Wonjune, et al.
Veröffentlicht: (2022)
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
Zero-Shot Text-to-Speech for Vietnamese
von: Vu, Thi, et al.
Veröffentlicht: (2025)
von: Vu, Thi, et al.
Veröffentlicht: (2025)
Transliterated Zero-Shot Domain Adaptation for Automatic Speech Recognition
von: Zhu, Han, et al.
Veröffentlicht: (2024)
von: Zhu, Han, et al.
Veröffentlicht: (2024)
Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
von: Zhou, Wangjin, et al.
Veröffentlicht: (2024)
von: Zhou, Wangjin, et al.
Veröffentlicht: (2024)
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning
von: Yang, Qian, et al.
Veröffentlicht: (2025)
von: Yang, Qian, et al.
Veröffentlicht: (2025)
Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning
von: He, Haorui, et al.
Veröffentlicht: (2024)
von: He, Haorui, et al.
Veröffentlicht: (2024)
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic Dialect Identification
von: Abdullah, Badr M., et al.
Veröffentlicht: (2025)
von: Abdullah, Badr M., et al.
Veröffentlicht: (2025)
Voice Conversion for Lombard Speaking Style with Implicit and Explicit Acoustic Feature Conditioning
von: Woszczyk, Dominika, et al.
Veröffentlicht: (2025)
von: Woszczyk, Dominika, et al.
Veröffentlicht: (2025)
An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems
von: Zhang, You, et al.
Veröffentlicht: (2021)
von: Zhang, You, et al.
Veröffentlicht: (2021)
Hard-Synth: Synthesizing Diverse Hard Samples for ASR using Zero-Shot TTS and LLM
von: Yu, Jiawei, et al.
Veröffentlicht: (2024)
von: Yu, Jiawei, et al.
Veröffentlicht: (2024)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
von: Doan, Khai Duy, et al.
Veröffentlicht: (2024)
von: Doan, Khai Duy, et al.
Veröffentlicht: (2024)
StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
von: Zhang, Yu, et al.
Veröffentlicht: (2023)
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
von: Dai, Shuqi, et al.
Veröffentlicht: (2025)
von: Dai, Shuqi, et al.
Veröffentlicht: (2025)
Marco-Voice Technical Report
von: Tian, Fengping, et al.
Veröffentlicht: (2025)
von: Tian, Fengping, et al.
Veröffentlicht: (2025)
LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
von: Chen, Shihao, et al.
Veröffentlicht: (2024)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion
von: Kamble, Anand, et al.
Veröffentlicht: (2023)
von: Kamble, Anand, et al.
Veröffentlicht: (2023)
ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2024)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2024)
Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech
von: Avdeeva, Anastasia, et al.
Veröffentlicht: (2024)
von: Avdeeva, Anastasia, et al.
Veröffentlicht: (2024)
Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model
von: Lehečka, Jan, et al.
Veröffentlicht: (2024)
von: Lehečka, Jan, et al.
Veröffentlicht: (2024)
Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost
von: Menta, Venkata Pushpak Teja
Veröffentlicht: (2026)
von: Menta, Venkata Pushpak Teja
Veröffentlicht: (2026)
Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed
von: Chen, Meiying, et al.
Veröffentlicht: (2022) -
SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
von: Li, Junjie, et al.
Veröffentlicht: (2023) -
PartialEdit: Identifying Partial Deepfakes in the Era of Neural Speech Editing
von: Zhang, You, et al.
Veröffentlicht: (2025) -
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
von: Cornell, Samuele, et al.
Veröffentlicht: (2024) -
Generating Novel and Realistic Speakers for Voice Conversion
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)