Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Jiang, Yuepeng, Li, Tao, Yang, Fengyu, Xie, Lei, Meng, Meng, Wang, Yujun |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
par: Lei, Shun, et autres
Publié: (2023)
par: Lei, Shun, et autres
Publié: (2023)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
par: Oh, Hyung-Seok, et autres
Publié: (2023)
par: Oh, Hyung-Seok, et autres
Publié: (2023)
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
par: Zhu, Xinfa, et autres
Publié: (2023)
par: Zhu, Xinfa, et autres
Publié: (2023)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
par: Wang, Tianrui, et autres
Publié: (2025)
par: Wang, Tianrui, et autres
Publié: (2025)
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
par: Ma, Guobin, et autres
Publié: (2025)
par: Ma, Guobin, et autres
Publié: (2025)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
par: Yang, Qian, et autres
Publié: (2024)
par: Yang, Qian, et autres
Publié: (2024)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
par: Wang, Helin, et autres
Publié: (2024)
par: Wang, Helin, et autres
Publié: (2024)
HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
par: Nishimura, Yuto, et autres
Publié: (2024)
par: Nishimura, Yuto, et autres
Publié: (2024)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
par: Wang, Chunhui, et autres
Publié: (2024)
par: Wang, Chunhui, et autres
Publié: (2024)
Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
par: Chen, Zhengyang, et autres
Publié: (2024)
par: Chen, Zhengyang, et autres
Publié: (2024)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
par: Jiang, Ziyue, et autres
Publié: (2023)
par: Jiang, Ziyue, et autres
Publié: (2023)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
par: Du, Chenpeng, et autres
Publié: (2021)
par: Du, Chenpeng, et autres
Publié: (2021)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
par: Chen, Weidong, et autres
Publié: (2025)
par: Chen, Weidong, et autres
Publié: (2025)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
par: Koriyama, Tomoki
Publié: (2025)
par: Koriyama, Tomoki
Publié: (2025)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
par: Li, Xuyuan, et autres
Publié: (2024)
par: Li, Xuyuan, et autres
Publié: (2024)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
par: Guan, Wenhao, et autres
Publié: (2023)
par: Guan, Wenhao, et autres
Publié: (2023)
SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
par: Guo, Haohan, et autres
Publié: (2024)
par: Guo, Haohan, et autres
Publié: (2024)
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
par: Zhang, Bowen, et autres
Publié: (2025)
par: Zhang, Bowen, et autres
Publié: (2025)
SoulX-Singer: Towards High-Quality Zero-Shot Singing Voice Synthesis
par: Qian, Jiale, et autres
Publié: (2026)
par: Qian, Jiale, et autres
Publié: (2026)
Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
par: Dai, Shuqi, et autres
Publié: (2025)
par: Dai, Shuqi, et autres
Publié: (2025)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
par: Ji, Shengpeng, et autres
Publié: (2024)
par: Ji, Shengpeng, et autres
Publié: (2024)
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
par: HU, Shujie, et autres
Publié: (2025)
par: HU, Shujie, et autres
Publié: (2025)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
par: Li, Yinghao Aaron, et autres
Publié: (2024)
par: Li, Yinghao Aaron, et autres
Publié: (2024)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
par: Yang, Dongchao, et autres
Publié: (2024)
par: Yang, Dongchao, et autres
Publié: (2024)
DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization
par: Chen, Huakang, et autres
Publié: (2025)
par: Chen, Huakang, et autres
Publié: (2025)
Hierarchical Control of Emotion Rendering in Speech Synthesis
par: Inoue, Sho, et autres
Publié: (2024)
par: Inoue, Sho, et autres
Publié: (2024)
From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models
par: Mu, Zhaoxi, et autres
Publié: (2025)
par: Mu, Zhaoxi, et autres
Publié: (2025)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
par: Wang, Yuanyuan, et autres
Publié: (2025)
par: Wang, Yuanyuan, et autres
Publié: (2025)
MultiVerse: Efficient and Expressive Zero-Shot Multi-Task Text-to-Speech
par: Bak, Taejun, et autres
Publié: (2024)
par: Bak, Taejun, et autres
Publié: (2024)
Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
par: Zhao, Junchuan, et autres
Publié: (2025)
par: Zhao, Junchuan, et autres
Publié: (2025)
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
par: Wang, Zhichao, et autres
Publié: (2024)
par: Wang, Zhichao, et autres
Publié: (2024)
Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling
par: Yang, Yuguang, et autres
Publié: (2024)
par: Yang, Yuguang, et autres
Publié: (2024)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
par: Inoue, Sho, et autres
Publié: (2024)
par: Inoue, Sho, et autres
Publié: (2024)
Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation
par: Han, Changjin, et autres
Publié: (2024)
par: Han, Changjin, et autres
Publié: (2024)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
par: Chen, Junyang, et autres
Publié: (2026)
par: Chen, Junyang, et autres
Publié: (2026)
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
par: Mahapatra, Aurosweta, et autres
Publié: (2025)
par: Mahapatra, Aurosweta, et autres
Publié: (2025)
Zero-Shot Mono-to-Binaural Speech Synthesis
par: Levkovitch, Alon, et autres
Publié: (2024)
par: Levkovitch, Alon, et autres
Publié: (2024)
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
par: Li, Haitao, et autres
Publié: (2026)
par: Li, Haitao, et autres
Publié: (2026)
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
par: Akti, Seymanur, et autres
Publié: (2025)
par: Akti, Seymanur, et autres
Publié: (2025)
Generative Expressive Conversational Speech Synthesis
par: Liu, Rui, et autres
Publié: (2024)
par: Liu, Rui, et autres
Publié: (2024)
Documents similaires
-
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
par: Lei, Shun, et autres
Publié: (2023) -
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
par: Oh, Hyung-Seok, et autres
Publié: (2023) -
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
par: Zhu, Xinfa, et autres
Publié: (2023) -
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
par: Wang, Tianrui, et autres
Publié: (2025) -
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
par: Ma, Guobin, et autres
Publié: (2025)