Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Du, Zongyang, Lu, Junchen, Zhou, Kun, Kaushik, Lakshmish, Sisman, Berrak |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
von: Du, Zongyang, et al.
Veröffentlicht: (2025)
von: Du, Zongyang, et al.
Veröffentlicht: (2025)
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025)
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
von: Lee, Philip H., et al.
Veröffentlicht: (2024)
von: Lee, Philip H., et al.
Veröffentlicht: (2024)
End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions
von: Kang, Wonjune, et al.
Veröffentlicht: (2022)
von: Kang, Wonjune, et al.
Veröffentlicht: (2022)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2026)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2026)
Style Mixture of Experts for Expressive Text-To-Speech Synthesis
von: Jawaid, Ahad, et al.
Veröffentlicht: (2024)
von: Jawaid, Ahad, et al.
Veröffentlicht: (2024)
Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion
von: Qi, Tianhua, et al.
Veröffentlicht: (2024)
von: Qi, Tianhua, et al.
Veröffentlicht: (2024)
Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2024)
Towards Naturalistic Voice Conversion: NaturalVoices Dataset with an Automatic Processing Pipeline
von: Salman, Ali N., et al.
Veröffentlicht: (2024)
von: Salman, Ali N., et al.
Veröffentlicht: (2024)
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
von: Guo, Zhao, et al.
Veröffentlicht: (2025)
von: Guo, Zhao, et al.
Veröffentlicht: (2025)
CONTUNER: Singing Voice Beautifying with Pitch and Expressiveness Condition
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2025)
von: Mahapatra, Aurosweta, et al.
Veröffentlicht: (2025)
LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2025)
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2025)
VISinger2+: End-to-End Singing Voice Synthesis Augmented by Self-Supervised Learning Representation
von: Yu, Yifeng, et al.
Veröffentlicht: (2024)
von: Yu, Yifeng, et al.
Veröffentlicht: (2024)
Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
von: Byun, Kyungguen, et al.
Veröffentlicht: (2025)
von: Byun, Kyungguen, et al.
Veröffentlicht: (2025)
End-to-End Integration of Speech Separation and Voice Activity Detection for Low-Latency Diarization of Telephone Conversations
von: Morrone, Giovanni, et al.
Veröffentlicht: (2023)
von: Morrone, Giovanni, et al.
Veröffentlicht: (2023)
Period Singer: Integrating Periodic and Aperiodic Variational Autoencoders for Natural-Sounding End-to-End Singing Voice Synthesis
von: Kim, Taewoo, et al.
Veröffentlicht: (2024)
von: Kim, Taewoo, et al.
Veröffentlicht: (2024)
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2024)
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2024)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
von: Melechovsky, Jan, et al.
Veröffentlicht: (2022)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2022)
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
von: Zhang, Qinglin, et al.
Veröffentlicht: (2024)
von: Zhang, Qinglin, et al.
Veröffentlicht: (2024)
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
von: Seki, Kentaro, et al.
Veröffentlicht: (2024)
von: Seki, Kentaro, et al.
Veröffentlicht: (2024)
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
von: Xu, Rixi, et al.
Veröffentlicht: (2026)
von: Xu, Rixi, et al.
Veröffentlicht: (2026)
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2026)
von: Wang, Zhichao, et al.
Veröffentlicht: (2026)
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models
von: Kim, Heeseung, et al.
Veröffentlicht: (2025)
von: Kim, Heeseung, et al.
Veröffentlicht: (2025)
JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis
von: Yu, Fan, et al.
Veröffentlicht: (2025)
von: Yu, Fan, et al.
Veröffentlicht: (2025)
Neural Concatenative Singing Voice Conversion: Rethinking Concatenation-Based Approach for One-Shot Singing Voice Conversion
von: Sha, Binzhu, et al.
Veröffentlicht: (2023)
von: Sha, Binzhu, et al.
Veröffentlicht: (2023)
VoiceGrad: Non-Parallel Any-to-Many Voice Conversion with Annealed Langevin Dynamics
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2020)
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2020)
CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion
von: Li, Yuke, et al.
Veröffentlicht: (2024)
von: Li, Yuke, et al.
Veröffentlicht: (2024)
ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
Enhancing Polyglot Voices by Leveraging Cross-Lingual Fine-Tuning in Any-to-One Voice Conversion
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2024)
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2024)
Generating Novel and Realistic Speakers for Voice Conversion
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
von: Akti, Seymanur, et al.
Veröffentlicht: (2025)
von: Akti, Seymanur, et al.
Veröffentlicht: (2025)
SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
von: Lam, Perry, et al.
Veröffentlicht: (2022)
von: Lam, Perry, et al.
Veröffentlicht: (2022)
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
von: Jung, Jaemin, et al.
Veröffentlicht: (2024)
von: Jung, Jaemin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
von: Du, Zongyang, et al.
Veröffentlicht: (2025) -
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
von: Ulgen, Ismail Rasim, et al.
Veröffentlicht: (2025) -
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
von: Lee, Philip H., et al.
Veröffentlicht: (2024) -
End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions
von: Kang, Wonjune, et al.
Veröffentlicht: (2022) -
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)