Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zuo, Jialong, Ji, Shengpeng, Fang, Minghui, Jiang, Ziyue, Cheng, Xize, Yang, Qian, Liu, Wenrui, Zhang, Guangyan, Tu, Zehai, Guo, Yiwen, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
CONTUNER: Singing Voice Beautifying with Pitch and Expressiveness Condition
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
von: Wang, Jianzong, et al.
Veröffentlicht: (2024)
Speech Watermarking with Discrete Intermediate Representations
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Enabling Beam Search for Language Model-Based Text-to-Speech Synthesis
von: Tu, Zehai, et al.
Veröffentlicht: (2024)
von: Tu, Zehai, et al.
Veröffentlicht: (2024)
TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
von: Ji, Shengpeng, et al.
Veröffentlicht: (2023)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2023)
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
von: Du, Zongyang, et al.
Veröffentlicht: (2024)
von: Du, Zongyang, et al.
Veröffentlicht: (2024)
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2025)
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2025)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed
von: Chen, Meiying, et al.
Veröffentlicht: (2022)
von: Chen, Meiying, et al.
Veröffentlicht: (2022)
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
von: Guan, Wenhao, et al.
Veröffentlicht: (2025)
von: Guan, Wenhao, et al.
Veröffentlicht: (2025)
Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
TSE-PI: Target Sound Extraction under Reverberant Environments with Pitch Information
von: Wang, Yiwen, et al.
Veröffentlicht: (2024)
von: Wang, Yiwen, et al.
Veröffentlicht: (2024)
Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching
von: Pan, Yu, et al.
Veröffentlicht: (2024)
von: Pan, Yu, et al.
Veröffentlicht: (2024)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching
von: Choi, Ha-Yeong, et al.
Veröffentlicht: (2025)
von: Choi, Ha-Yeong, et al.
Veröffentlicht: (2025)
SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
von: Bai, Bingsong, et al.
Veröffentlicht: (2024)
von: Bai, Bingsong, et al.
Veröffentlicht: (2024)
WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
von: Ji, Shengpeng, et al.
Veröffentlicht: (2025)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2025)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
Adversarial Multi-Task Learning for Disentangling Timbre and Pitch in Singing Voice Synthesis
von: Kim, Tae-Woo, et al.
Veröffentlicht: (2022)
von: Kim, Tae-Woo, et al.
Veröffentlicht: (2022)
Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
Discrete Optimal Transport and Voice Conversion
von: Selitskiy, Anton, et al.
Veröffentlicht: (2025)
von: Selitskiy, Anton, et al.
Veröffentlicht: (2025)
VoiceRestore: Flow-Matching Transformers for Speech Recording Quality Restoration
von: Kirdey, Stanislav
Veröffentlicht: (2025)
von: Kirdey, Stanislav
Veröffentlicht: (2025)
Entropy-based Coarse and Compressed Semantic Speech Representation Learning
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis
von: Yu, Fan, et al.
Veröffentlicht: (2025)
von: Yu, Fan, et al.
Veröffentlicht: (2025)
DJCM: A Deep Joint Cascade Model for Singing Voice Separation and Vocal Pitch Estimation
von: Wei, Haojie, et al.
Veröffentlicht: (2024)
von: Wei, Haojie, et al.
Veröffentlicht: (2024)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
von: Seki, Kentaro, et al.
Veröffentlicht: (2024)
von: Seki, Kentaro, et al.
Veröffentlicht: (2024)
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
von: Akti, Seymanur, et al.
Veröffentlicht: (2025)
von: Akti, Seymanur, et al.
Veröffentlicht: (2025)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
von: Zhu, Han, et al.
Veröffentlicht: (2025)
von: Zhu, Han, et al.
Veröffentlicht: (2025)
FlowW2N: Whispered-to-Normal Speech Conversion via Flow-Matching
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2026)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2026)
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
von: Igarashi, Takuto, et al.
Veröffentlicht: (2024)
von: Igarashi, Takuto, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
von: Zuo, Jialong, et al.
Veröffentlicht: (2025) -
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024) -
CONTUNER: Singing Voice Beautifying with Pitch and Expressiveness Condition
von: Wang, Jianzong, et al.
Veröffentlicht: (2024) -
Speech Watermarking with Discrete Intermediate Representations
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024) -
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)