DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Weidong, Yang, Shan, Li, Guangzhi, Wu, Xixin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
di: Yang, Qian, et al.
Pubblicazione: (2024)
di: Yang, Qian, et al.
Pubblicazione: (2024)
Prosodic Parameter Manipulation in TTS generated speech for Controlled Speech Generation
di: Chary, Podakanti Satyajith
Pubblicazione: (2024)
di: Chary, Podakanti Satyajith
Pubblicazione: (2024)
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
di: Li, Zhu, et al.
Pubblicazione: (2025)
di: Li, Zhu, et al.
Pubblicazione: (2025)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
di: Yang, Dongchao, et al.
Pubblicazione: (2024)
di: Yang, Dongchao, et al.
Pubblicazione: (2024)
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
di: Chen, Xueyuan, et al.
Pubblicazione: (2025)
di: Chen, Xueyuan, et al.
Pubblicazione: (2025)
DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis
di: Gu, Yu, et al.
Pubblicazione: (2024)
di: Gu, Yu, et al.
Pubblicazione: (2024)
SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
di: Guo, Haohan, et al.
Pubblicazione: (2024)
di: Guo, Haohan, et al.
Pubblicazione: (2024)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
di: Jiang, Yuepeng, et al.
Pubblicazione: (2024)
di: Jiang, Yuepeng, et al.
Pubblicazione: (2024)
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
di: Wu, Wenxuan, et al.
Pubblicazione: (2024)
di: Wu, Wenxuan, et al.
Pubblicazione: (2024)
PSST! Prosodic Speech Segmentation with Transformers
di: Roll, Nathan, et al.
Pubblicazione: (2023)
di: Roll, Nathan, et al.
Pubblicazione: (2023)
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
di: Li, Hanzhao, et al.
Pubblicazione: (2025)
di: Li, Hanzhao, et al.
Pubblicazione: (2025)
Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
di: Ohnaka, Hien, et al.
Pubblicazione: (2025)
di: Ohnaka, Hien, et al.
Pubblicazione: (2025)
UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization
di: Wang, Yuejiao, et al.
Pubblicazione: (2024)
di: Wang, Yuejiao, et al.
Pubblicazione: (2024)
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
di: Zhu, Xinfa, et al.
Pubblicazione: (2023)
di: Zhu, Xinfa, et al.
Pubblicazione: (2023)
Generative Expressive Conversational Speech Synthesis
di: Liu, Rui, et al.
Pubblicazione: (2024)
di: Liu, Rui, et al.
Pubblicazione: (2024)
Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
di: Guo, Haohan, et al.
Pubblicazione: (2024)
di: Guo, Haohan, et al.
Pubblicazione: (2024)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
di: Guan, Wenhao, et al.
Pubblicazione: (2023)
di: Guan, Wenhao, et al.
Pubblicazione: (2023)
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis
di: Niu, Rui, et al.
Pubblicazione: (2025)
di: Niu, Rui, et al.
Pubblicazione: (2025)
Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
di: Onda, Kentaro, et al.
Pubblicazione: (2025)
di: Onda, Kentaro, et al.
Pubblicazione: (2025)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
di: Lei, Shun, et al.
Pubblicazione: (2023)
di: Lei, Shun, et al.
Pubblicazione: (2023)
UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
di: Wang, Yuanyuan, et al.
Pubblicazione: (2026)
di: Wang, Yuanyuan, et al.
Pubblicazione: (2026)
EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
di: de Seyssel, Maureen, et al.
Pubblicazione: (2023)
di: de Seyssel, Maureen, et al.
Pubblicazione: (2023)
UniSep: Universal Target Audio Separation with Language Models at Scale
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
di: Deng, Yimin, et al.
Pubblicazione: (2024)
di: Deng, Yimin, et al.
Pubblicazione: (2024)
Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder
di: Guo, Haohan, et al.
Pubblicazione: (2024)
di: Guo, Haohan, et al.
Pubblicazione: (2024)
Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
di: Chen, Xueyuan, et al.
Pubblicazione: (2024)
di: Chen, Xueyuan, et al.
Pubblicazione: (2024)
Hierarchical Control of Emotion Rendering in Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
di: Xie, Jingran, et al.
Pubblicazione: (2025)
di: Xie, Jingran, et al.
Pubblicazione: (2025)
DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models
di: wu, Weihao, et al.
Pubblicazione: (2025)
di: wu, Weihao, et al.
Pubblicazione: (2025)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
Cross-Utterance Conditioned VAE for Speech Generation
di: Li, Yang, et al.
Pubblicazione: (2023)
di: Li, Yang, et al.
Pubblicazione: (2023)
NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation
di: Ni, Qinke, et al.
Pubblicazione: (2026)
di: Ni, Qinke, et al.
Pubblicazione: (2026)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
di: Chen, Peikun, et al.
Pubblicazione: (2024)
di: Chen, Peikun, et al.
Pubblicazione: (2024)
Autoregressive Speech Synthesis without Vector Quantization
di: Meng, Lingwei, et al.
Pubblicazione: (2024)
di: Meng, Lingwei, et al.
Pubblicazione: (2024)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
di: Gong, Cheng, et al.
Pubblicazione: (2023)
di: Gong, Cheng, et al.
Pubblicazione: (2023)
ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
di: Lou, Haowei, et al.
Pubblicazione: (2025)
di: Lou, Haowei, et al.
Pubblicazione: (2025)
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
di: Deng, Keqi, et al.
Pubblicazione: (2024)
di: Deng, Keqi, et al.
Pubblicazione: (2024)
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
di: Tao, Dehua, et al.
Pubblicazione: (2024)
di: Tao, Dehua, et al.
Pubblicazione: (2024)
VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
di: Zhou, Yixuan, et al.
Pubblicazione: (2024)
di: Zhou, Yixuan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
di: Yang, Qian, et al.
Pubblicazione: (2024) -
Prosodic Parameter Manipulation in TTS generated speech for Controlled Speech Generation
di: Chary, Podakanti Satyajith
Pubblicazione: (2024) -
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
di: Li, Zhu, et al.
Pubblicazione: (2025) -
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
di: Yang, Dongchao, et al.
Pubblicazione: (2024) -
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
di: Chen, Xueyuan, et al.
Pubblicazione: (2025)