DMP-TTS: Disentangled multi-modal Prompting for Controllable Text-to-Speech with Chained Guidance
Fuente:
arXiv
Guardado en:
| Autores principales: | Yin, Kang, Qiang, Chunyu, Zhao, Sirui, Wang, Xiaopeng, Liang, Yuzhe, Cai, Pengfei, Xu, Tong, Zhang, Chen, Chen, Enhong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis
por: Wang, Xiaopeng, et al.
Publicado: (2025)
por: Wang, Xiaopeng, et al.
Publicado: (2025)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
por: Guan, Wenhao, et al.
Publicado: (2023)
por: Guan, Wenhao, et al.
Publicado: (2023)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
por: Fu, Ruibo, et al.
Publicado: (2024)
por: Fu, Ruibo, et al.
Publicado: (2024)
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
por: Qiang, Chunyu, et al.
Publicado: (2026)
por: Qiang, Chunyu, et al.
Publicado: (2026)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
por: Lu, Ye-Xin, et al.
Publicado: (2025)
por: Lu, Ye-Xin, et al.
Publicado: (2025)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
por: Jiang, Ziyue, et al.
Publicado: (2023)
por: Jiang, Ziyue, et al.
Publicado: (2023)
TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis
por: Liang, Qifan, et al.
Publicado: (2026)
por: Liang, Qifan, et al.
Publicado: (2026)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
por: Han, Wooseok, et al.
Publicado: (2024)
por: Han, Wooseok, et al.
Publicado: (2024)
DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
por: Qi, Xin, et al.
Publicado: (2024)
por: Qi, Xin, et al.
Publicado: (2024)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
por: Wu, Zhichao, et al.
Publicado: (2025)
por: Wu, Zhichao, et al.
Publicado: (2025)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
por: Liu, Huadai, et al.
Publicado: (2023)
por: Liu, Huadai, et al.
Publicado: (2023)
Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis
por: Chen, Yushen, et al.
Publicado: (2026)
por: Chen, Yushen, et al.
Publicado: (2026)
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS
por: Wang, Haoyu, et al.
Publicado: (2024)
por: Wang, Haoyu, et al.
Publicado: (2024)
MM-Sonate: Multimodal Controllable Audio-Video Generation with Zero-Shot Voice Cloning
por: Qiang, Chunyu, et al.
Publicado: (2026)
por: Qiang, Chunyu, et al.
Publicado: (2026)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
por: Ma, Ziyang, et al.
Publicado: (2023)
por: Ma, Ziyang, et al.
Publicado: (2023)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
por: Guo, Hao-Han, et al.
Publicado: (2025)
por: Guo, Hao-Han, et al.
Publicado: (2025)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
por: Wang, Chunhui, et al.
Publicado: (2024)
por: Wang, Chunhui, et al.
Publicado: (2024)
Improved Dysarthric Speech to Text Conversion via TTS Personalization
por: Mihajlik, Péter, et al.
Publicado: (2025)
por: Mihajlik, Péter, et al.
Publicado: (2025)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
por: Guo, Hao-Han, et al.
Publicado: (2024)
por: Guo, Hao-Han, et al.
Publicado: (2024)
Probing Human Articulatory Constraints in End-to-End TTS with Reverse and Mismatched Speech-Text Directions
por: Khadse, Parth, et al.
Publicado: (2026)
por: Khadse, Parth, et al.
Publicado: (2026)
EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech
por: Liang, Ziqi, et al.
Publicado: (2024)
por: Liang, Ziqi, et al.
Publicado: (2024)
AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis
por: Luo, Dan, et al.
Publicado: (2025)
por: Luo, Dan, et al.
Publicado: (2025)
MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control
por: Mai, Jialong, et al.
Publicado: (2026)
por: Mai, Jialong, et al.
Publicado: (2026)
MunTTS: A Text-to-Speech System for Mundari
por: Gumma, Varun, et al.
Publicado: (2024)
por: Gumma, Varun, et al.
Publicado: (2024)
InstructAudio: Unified speech and music generation with natural language instruction
por: Qiang, Chunyu, et al.
Publicado: (2025)
por: Qiang, Chunyu, et al.
Publicado: (2025)
Communication-Efficient Personalized Federated Learning for Speech-to-Text Tasks
por: Du, Yichao, et al.
Publicado: (2024)
por: Du, Yichao, et al.
Publicado: (2024)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
por: Wang, Xinsheng, et al.
Publicado: (2025)
por: Wang, Xinsheng, et al.
Publicado: (2025)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
por: Ren, Yong, et al.
Publicado: (2026)
por: Ren, Yong, et al.
Publicado: (2026)
ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
por: Lou, Haowei, et al.
Publicado: (2025)
por: Lou, Haowei, et al.
Publicado: (2025)
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
por: Deng, Wei, et al.
Publicado: (2025)
por: Deng, Wei, et al.
Publicado: (2025)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
por: Cho, Deok-Hyeon, et al.
Publicado: (2025)
por: Cho, Deok-Hyeon, et al.
Publicado: (2025)
SynTTS-Commands: A Public Dataset for On-Device KWS via TTS-Synthesized Multilingual Speech
por: Gan, Lu, et al.
Publicado: (2025)
por: Gan, Lu, et al.
Publicado: (2025)
PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation
por: Shi, Shuchen, et al.
Publicado: (2024)
por: Shi, Shuchen, et al.
Publicado: (2024)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
por: Wang, Tianrui, et al.
Publicado: (2025)
por: Wang, Tianrui, et al.
Publicado: (2025)
Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis
por: Liu, Qingyu, et al.
Publicado: (2025)
por: Liu, Qingyu, et al.
Publicado: (2025)
Prototype-Based Disentanglement for Controllable Dysarthric Speech Synthesis
por: Wang, Haoshen, et al.
Publicado: (2026)
por: Wang, Haoshen, et al.
Publicado: (2026)
StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations
por: Liu, Sen, et al.
Publicado: (2024)
por: Liu, Sen, et al.
Publicado: (2024)
RephraseTTS: Dynamic Length Text based Speech Insertion with Speaker Style Transfer
por: Matiyali, Neeraj, et al.
Publicado: (2025)
por: Matiyali, Neeraj, et al.
Publicado: (2025)
LoRP-TTS: Low-Rank Personalized Text-To-Speech
por: Bondaruk, Łukasz, et al.
Publicado: (2025)
por: Bondaruk, Łukasz, et al.
Publicado: (2025)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
por: Xue, Heyang, et al.
Publicado: (2025)
por: Xue, Heyang, et al.
Publicado: (2025)
Ejemplares similares
-
M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis
por: Wang, Xiaopeng, et al.
Publicado: (2025) -
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
por: Guan, Wenhao, et al.
Publicado: (2023) -
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
por: Fu, Ruibo, et al.
Publicado: (2024) -
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
por: Qiang, Chunyu, et al.
Publicado: (2026) -
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
por: Lu, Ye-Xin, et al.
Publicado: (2025)