TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Liang, Qifan, Liu, Yuansen, Wei, Ruixin, Lu, Nan, Zhao, Junchuan, Wang, Ye |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control
di: Mai, Jialong, et al.
Pubblicazione: (2026)
di: Mai, Jialong, et al.
Pubblicazione: (2026)
EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering
di: Xie, Tianxin, et al.
Pubblicazione: (2025)
di: Xie, Tianxin, et al.
Pubblicazione: (2025)
Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection
di: Zhao, Junchuan, et al.
Pubblicazione: (2026)
di: Zhao, Junchuan, et al.
Pubblicazione: (2026)
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
di: Zhou, Siyi, et al.
Pubblicazione: (2025)
di: Zhou, Siyi, et al.
Pubblicazione: (2025)
CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance
di: Zhao, Junchuan, et al.
Pubblicazione: (2025)
di: Zhao, Junchuan, et al.
Pubblicazione: (2025)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
di: Lu, Ye-Xin, et al.
Pubblicazione: (2025)
di: Lu, Ye-Xin, et al.
Pubblicazione: (2025)
EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech
di: Liang, Ziqi, et al.
Pubblicazione: (2024)
di: Liang, Ziqi, et al.
Pubblicazione: (2024)
DMP-TTS: Disentangled multi-modal Prompting for Controllable Text-to-Speech with Chained Guidance
di: Yin, Kang, et al.
Pubblicazione: (2025)
di: Yin, Kang, et al.
Pubblicazione: (2025)
Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis
di: Feng, Pengchao, et al.
Pubblicazione: (2025)
di: Feng, Pengchao, et al.
Pubblicazione: (2025)
Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification
di: Zhao, Yiyang, et al.
Pubblicazione: (2025)
di: Zhao, Yiyang, et al.
Pubblicazione: (2025)
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
di: Li, Haoxun, et al.
Pubblicazione: (2025)
di: Li, Haoxun, et al.
Pubblicazione: (2025)
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
di: Deng, Wei, et al.
Pubblicazione: (2025)
di: Deng, Wei, et al.
Pubblicazione: (2025)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
EMORL-TTS: Reinforcement Learning for Fine-Grained Emotion Control in LLM-based TTS
di: Li, Haoxun, et al.
Pubblicazione: (2025)
di: Li, Haoxun, et al.
Pubblicazione: (2025)
ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
di: Tang, Haobin, et al.
Pubblicazione: (2024)
di: Tang, Haobin, et al.
Pubblicazione: (2024)
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2024)
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2024)
Multi-Utterance Speech Separation and Association Trained on Short Segments
di: Wang, Yuzhu, et al.
Pubblicazione: (2025)
di: Wang, Yuzhu, et al.
Pubblicazione: (2025)
AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis
di: Luo, Dan, et al.
Pubblicazione: (2025)
di: Luo, Dan, et al.
Pubblicazione: (2025)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
di: Jiang, Ziyue, et al.
Pubblicazione: (2023)
di: Jiang, Ziyue, et al.
Pubblicazione: (2023)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
di: Gudmalwar, Ashishkumar, et al.
Pubblicazione: (2024)
di: Gudmalwar, Ashishkumar, et al.
Pubblicazione: (2024)
NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech
di: Borisov, Maksim, et al.
Pubblicazione: (2025)
di: Borisov, Maksim, et al.
Pubblicazione: (2025)
VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
di: Zhou, Yixuan, et al.
Pubblicazione: (2025)
di: Zhou, Yixuan, et al.
Pubblicazione: (2025)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
di: Liu, Huadai, et al.
Pubblicazione: (2023)
di: Liu, Huadai, et al.
Pubblicazione: (2023)
Disentangling Score Content and Performance Style for Joint Piano Rendering and Transcription
di: Zeng, Wei, et al.
Pubblicazione: (2025)
di: Zeng, Wei, et al.
Pubblicazione: (2025)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
di: Guo, Yinlin, et al.
Pubblicazione: (2024)
di: Guo, Yinlin, et al.
Pubblicazione: (2024)
Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis
di: Liu, Qingyu, et al.
Pubblicazione: (2025)
di: Liu, Qingyu, et al.
Pubblicazione: (2025)
SynTTS-Commands: A Public Dataset for On-Device KWS via TTS-Synthesized Multilingual Speech
di: Gan, Lu, et al.
Pubblicazione: (2025)
di: Gan, Lu, et al.
Pubblicazione: (2025)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2025)
di: Inoue, Sho, et al.
Pubblicazione: (2025)
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs
di: Zhang, Enshi, et al.
Pubblicazione: (2024)
di: Zhang, Enshi, et al.
Pubblicazione: (2024)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
di: Fu, Ruibo, et al.
Pubblicazione: (2024)
di: Fu, Ruibo, et al.
Pubblicazione: (2024)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
di: Ren, Yong, et al.
Pubblicazione: (2026)
di: Ren, Yong, et al.
Pubblicazione: (2026)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
di: Guan, Wenhao, et al.
Pubblicazione: (2023)
di: Guan, Wenhao, et al.
Pubblicazione: (2023)
Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
di: Zhao, Junchuan, et al.
Pubblicazione: (2025)
di: Zhao, Junchuan, et al.
Pubblicazione: (2025)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
di: Wang, Xinsheng, et al.
Pubblicazione: (2025)
di: Wang, Xinsheng, et al.
Pubblicazione: (2025)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
di: Peng, Puyuan, et al.
Pubblicazione: (2025)
di: Peng, Puyuan, et al.
Pubblicazione: (2025)
M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis
di: Wang, Xiaopeng, et al.
Pubblicazione: (2025)
di: Wang, Xiaopeng, et al.
Pubblicazione: (2025)
LLaDA-TTS: Unifying Speech Synthesis and Zero-Shot Editing via Masked Diffusion Modeling
di: Fan, Xiaoyu, et al.
Pubblicazione: (2026)
di: Fan, Xiaoyu, et al.
Pubblicazione: (2026)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
di: Wu, Zhichao, et al.
Pubblicazione: (2025)
di: Wu, Zhichao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control
di: Mai, Jialong, et al.
Pubblicazione: (2026) -
EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering
di: Xie, Tianxin, et al.
Pubblicazione: (2025) -
Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection
di: Zhao, Junchuan, et al.
Pubblicazione: (2026) -
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
di: Zhou, Siyi, et al.
Pubblicazione: (2025) -
CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance
di: Zhao, Junchuan, et al.
Pubblicazione: (2025)