Total-Duration-Aware Duration Modeling for Text-to-Speech Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Eskimez, Sefik Emre, Wang, Xiaofei, Thakker, Manthan, Tsai, Chung-Hsien, Li, Canrun, Xiao, Zhen, Yang, Hemin, Zhu, Zirun, Tang, Min, Li, Jinyu, Zhao, Sheng, Kanda, Naoyuki |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
di: Eskimez, Sefik Emre, et al.
Pubblicazione: (2024)
di: Eskimez, Sefik Emre, et al.
Pubblicazione: (2024)
Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like
di: Kanda, Naoyuki, et al.
Pubblicazione: (2024)
di: Kanda, Naoyuki, et al.
Pubblicazione: (2024)
An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS
di: Wang, Xiaofei, et al.
Pubblicazione: (2024)
di: Wang, Xiaofei, et al.
Pubblicazione: (2024)
TS3-Codec: Transformer-Based Simple Streaming Single Codec
di: Wu, Haibin, et al.
Pubblicazione: (2024)
di: Wu, Haibin, et al.
Pubblicazione: (2024)
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
di: Wang, Xiaofei, et al.
Pubblicazione: (2023)
di: Wang, Xiaofei, et al.
Pubblicazione: (2023)
Neural Speech Extraction with Human Feedback
di: Itani, Malek, et al.
Pubblicazione: (2025)
di: Itani, Malek, et al.
Pubblicazione: (2025)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
di: Li, Jiaqi, et al.
Pubblicazione: (2024)
di: Li, Jiaqi, et al.
Pubblicazione: (2024)
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
di: Subramanian, Aswin Shanmugam, et al.
Pubblicazione: (2025)
di: Subramanian, Aswin Shanmugam, et al.
Pubblicazione: (2025)
Adaptive Duration Model for Text Speech Alignment
di: Cao, Junjie
Pubblicazione: (2025)
di: Cao, Junjie
Pubblicazione: (2025)
DiariST: Streaming Speech Translation with Speaker Diarization
di: Yang, Mu, et al.
Pubblicazione: (2023)
di: Yang, Mu, et al.
Pubblicazione: (2023)
DAME: Duration-Aware Matryoshka Embedding for Duration-Robust Speaker Verification
di: Jung, Youngmoon, et al.
Pubblicazione: (2026)
di: Jung, Youngmoon, et al.
Pubblicazione: (2026)
FNH-TTS: Mixture-of-Experts Duration Modeling for Robust Neural Speech Synthesis
di: Meng, Qingliang, et al.
Pubblicazione: (2025)
di: Meng, Qingliang, et al.
Pubblicazione: (2025)
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis
di: Li, Yinghao Aaron, et al.
Pubblicazione: (2025)
di: Li, Yinghao Aaron, et al.
Pubblicazione: (2025)
DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing
di: Sahipjohn, Neha, et al.
Pubblicazione: (2024)
di: Sahipjohn, Neha, et al.
Pubblicazione: (2024)
Knowledge boosting during low-latency inference
di: Srinivas, Vidya, et al.
Pubblicazione: (2024)
di: Srinivas, Vidya, et al.
Pubblicazione: (2024)
The Overview of Segmental Durations Modification Algorithms on Speech Signal Characteristics
di: Jang, Kyeomeun, et al.
Pubblicazione: (2025)
di: Jang, Kyeomeun, et al.
Pubblicazione: (2025)
Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration Modeling
di: Prabhu, Navin Raj, et al.
Pubblicazione: (2025)
di: Prabhu, Navin Raj, et al.
Pubblicazione: (2025)
Target conversation extraction: Source separation using turn-taking dynamics
di: Chen, Tuochao, et al.
Pubblicazione: (2024)
di: Chen, Tuochao, et al.
Pubblicazione: (2024)
CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching
di: Zhang, Leying, et al.
Pubblicazione: (2025)
di: Zhang, Leying, et al.
Pubblicazione: (2025)
Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding
di: Zheng, Rui-Chen, et al.
Pubblicazione: (2025)
di: Zheng, Rui-Chen, et al.
Pubblicazione: (2025)
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis
di: Lu, Ye-Xin, et al.
Pubblicazione: (2024)
di: Lu, Ye-Xin, et al.
Pubblicazione: (2024)
DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis
di: Gu, Yu, et al.
Pubblicazione: (2024)
di: Gu, Yu, et al.
Pubblicazione: (2024)
Position: Towards Responsible Evaluation for Text-to-Speech
di: Yang, Yifan, et al.
Pubblicazione: (2025)
di: Yang, Yifan, et al.
Pubblicazione: (2025)
Conversational Speech Naturalness Predictor
di: Xu, Anfeng, et al.
Pubblicazione: (2026)
di: Xu, Anfeng, et al.
Pubblicazione: (2026)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
di: Peng, Puyuan, et al.
Pubblicazione: (2025)
di: Peng, Puyuan, et al.
Pubblicazione: (2025)
Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
di: Wang, Peidong, et al.
Pubblicazione: (2025)
di: Wang, Peidong, et al.
Pubblicazione: (2025)
Phone Duration Modeling for Speaker Age Estimation in Children
di: Shivakumar, Prashanth Gurunath, et al.
Pubblicazione: (2021)
di: Shivakumar, Prashanth Gurunath, et al.
Pubblicazione: (2021)
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
di: Yang, Yifan, et al.
Pubblicazione: (2024)
di: Yang, Yifan, et al.
Pubblicazione: (2024)
Towards Efficient Speech-Text Jointly Decoding within One Speech Language Model
di: Wu, Haibin, et al.
Pubblicazione: (2025)
di: Wu, Haibin, et al.
Pubblicazione: (2025)
DASS: Distilled Audio State Space Models Are Stronger and More Duration-Scalable Learners
di: Bhati, Saurabhchand, et al.
Pubblicazione: (2024)
di: Bhati, Saurabhchand, et al.
Pubblicazione: (2024)
Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration
di: Lou, Haowei, et al.
Pubblicazione: (2024)
di: Lou, Haowei, et al.
Pubblicazione: (2024)
A Composite Predictive-Generative Approach to Monaural Universal Speech Enhancement
di: Zhang, Jie, et al.
Pubblicazione: (2025)
di: Zhang, Jie, et al.
Pubblicazione: (2025)
Spoken Language Modeling with Duration-Penalized Self-Supervised Units
di: Visser, Nicol, et al.
Pubblicazione: (2025)
di: Visser, Nicol, et al.
Pubblicazione: (2025)
Multichannel Long-Term Streaming Neural Speech Enhancement for Static and Moving Speakers
di: Quan, Changsheng, et al.
Pubblicazione: (2024)
di: Quan, Changsheng, et al.
Pubblicazione: (2024)
Mel-FullSubNet: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR
di: Zhou, Rui, et al.
Pubblicazione: (2024)
di: Zhou, Rui, et al.
Pubblicazione: (2024)
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
di: Zhou, Siyi, et al.
Pubblicazione: (2025)
di: Zhou, Siyi, et al.
Pubblicazione: (2025)
Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation
di: Lou, Haowei, et al.
Pubblicazione: (2025)
di: Lou, Haowei, et al.
Pubblicazione: (2025)
ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
di: Chen, Yafeng, et al.
Pubblicazione: (2024)
di: Chen, Yafeng, et al.
Pubblicazione: (2024)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
di: Lu, Ye-Xin, et al.
Pubblicazione: (2025)
di: Lu, Ye-Xin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech
di: Wu, Haibin, et al.
Pubblicazione: (2024) -
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
di: Eskimez, Sefik Emre, et al.
Pubblicazione: (2024) -
Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like
di: Kanda, Naoyuki, et al.
Pubblicazione: (2024) -
An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS
di: Wang, Xiaofei, et al.
Pubblicazione: (2024) -
TS3-Codec: Transformer-Based Simple Streaming Single Codec
di: Wu, Haibin, et al.
Pubblicazione: (2024)