Towards Developing State-of-the-Art TTS Synthesisers for 13 Indian Languages with Signal Processing aided Alignments
Fuente:
arXiv
Guardado en:
| Autores principales: | Prakash, Anusha, Umesh, S, Murthy, Hema A |
|---|---|
| Formato: | Preprint |
| Publicado: |
2022
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Exploring an Inter-Pausal Unit (IPU) based Approach for Indic End-to-End TTS Systems
por: Prakash, Anusha, et al.
Publicado: (2024)
por: Prakash, Anusha, et al.
Publicado: (2024)
AutoProsody: A Prosodic Feature Extraction Tool for Indian Languages
por: Thinakaran, Preethi, et al.
Publicado: (2025)
por: Thinakaran, Preethi, et al.
Publicado: (2025)
A Unified Framework for Collecting Text-to-Speech Synthesis Datasets for 22 Indian Languages
por: Sathiyamoorthy, Sujitha, et al.
Publicado: (2024)
por: Sathiyamoorthy, Sujitha, et al.
Publicado: (2024)
Intelli-Z: Toward Intelligible Zero-Shot TTS
por: Jung, Sunghee, et al.
Publicado: (2024)
por: Jung, Sunghee, et al.
Publicado: (2024)
SponTTS: modeling and transferring spontaneous style for TTS
por: Li, Hanzhao, et al.
Publicado: (2023)
por: Li, Hanzhao, et al.
Publicado: (2023)
E1 TTS: Simple and Fast Non-Autoregressive TTS
por: Liu, Zhijun, et al.
Publicado: (2024)
por: Liu, Zhijun, et al.
Publicado: (2024)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
por: Ren, Yong, et al.
Publicado: (2026)
por: Ren, Yong, et al.
Publicado: (2026)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
por: Eskimez, Sefik Emre, et al.
Publicado: (2024)
por: Eskimez, Sefik Emre, et al.
Publicado: (2024)
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
por: Qharabagh, Mahta Fetrat, et al.
Publicado: (2024)
por: Qharabagh, Mahta Fetrat, et al.
Publicado: (2024)
A2TTS: TTS for Low Resource Indian Languages
por: Bhadoriya, Ayush Singh, et al.
Publicado: (2025)
por: Bhadoriya, Ayush Singh, et al.
Publicado: (2025)
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
por: Xie, Kun, et al.
Publicado: (2025)
por: Xie, Kun, et al.
Publicado: (2025)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
por: Xue, Heyang, et al.
Publicado: (2025)
por: Xue, Heyang, et al.
Publicado: (2025)
Combining Masked Language Modeling and Cross-Modal Contrastive Learning for Prosody-Aware TTS
por: Borodin, Kirill, et al.
Publicado: (2026)
por: Borodin, Kirill, et al.
Publicado: (2026)
ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
por: Lou, Haowei, et al.
Publicado: (2025)
por: Lou, Haowei, et al.
Publicado: (2025)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
por: Kim, Jaehyeon, et al.
Publicado: (2024)
por: Kim, Jaehyeon, et al.
Publicado: (2024)
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
por: Guan, Wenhao, et al.
Publicado: (2025)
por: Guan, Wenhao, et al.
Publicado: (2025)
Accent-VITS:accent transfer for end-to-end TTS
por: Ma, Linhan, et al.
Publicado: (2023)
por: Ma, Linhan, et al.
Publicado: (2023)
A Dataset for Automatic Assessment of TTS Quality in Spanish
por: Welford, Alejandro Sosa, et al.
Publicado: (2025)
por: Welford, Alejandro Sosa, et al.
Publicado: (2025)
Zero-shot Cross-lingual Voice Transfer for TTS
por: Biadsy, Fadi, et al.
Publicado: (2024)
por: Biadsy, Fadi, et al.
Publicado: (2024)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
por: Liu, Huadai, et al.
Publicado: (2023)
por: Liu, Huadai, et al.
Publicado: (2023)
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
por: Nguyen, Tan Dat, et al.
Publicado: (2025)
por: Nguyen, Tan Dat, et al.
Publicado: (2025)
Enhancing TTS Stability in Hebrew using Discrete Semantic Units
por: Zeldes, Ella, et al.
Publicado: (2024)
por: Zeldes, Ella, et al.
Publicado: (2024)
Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
por: He, Xinlu, et al.
Publicado: (2025)
por: He, Xinlu, et al.
Publicado: (2025)
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
por: Li, Haoxun, et al.
Publicado: (2025)
por: Li, Haoxun, et al.
Publicado: (2025)
SPAM: Style Prompt Adherence Metric for Prompt-based TTS
por: Cho, Chanhee, et al.
Publicado: (2026)
por: Cho, Chanhee, et al.
Publicado: (2026)
Low-Resource Self-Supervised Learning with SSL-Enhanced TTS
por: Hsu, Po-chun, et al.
Publicado: (2023)
por: Hsu, Po-chun, et al.
Publicado: (2023)
The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024
por: Zhou, Shuoyi, et al.
Publicado: (2024)
por: Zhou, Shuoyi, et al.
Publicado: (2024)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
por: Jeon, Yejin, et al.
Publicado: (2024)
por: Jeon, Yejin, et al.
Publicado: (2024)
Bridging the gap between training and inference in LM-based TTS models
por: Zhang, Ruonan, et al.
Publicado: (2025)
por: Zhang, Ruonan, et al.
Publicado: (2025)
Prosodic Parameter Manipulation in TTS generated speech for Controlled Speech Generation
por: Chary, Podakanti Satyajith
Publicado: (2024)
por: Chary, Podakanti Satyajith
Publicado: (2024)
On the relationship between speech and hearing
por: Umesh, Srinivasan, et al.
Publicado: (2024)
por: Umesh, Srinivasan, et al.
Publicado: (2024)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
por: Jiang, Ziyue, et al.
Publicado: (2023)
por: Jiang, Ziyue, et al.
Publicado: (2023)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
por: Guo, Hao-Han, et al.
Publicado: (2025)
por: Guo, Hao-Han, et al.
Publicado: (2025)
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
por: Anastassiou, Philip, et al.
Publicado: (2024)
por: Anastassiou, Philip, et al.
Publicado: (2024)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
por: Peng, Puyuan, et al.
Publicado: (2025)
por: Peng, Puyuan, et al.
Publicado: (2025)
DQR-TTS: Semi-supervised Text-to-speech Synthesis with Dynamic Quantized Representation
por: Wang, Jianzong, et al.
Publicado: (2023)
por: Wang, Jianzong, et al.
Publicado: (2023)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
por: Lu, Ye-Xin, et al.
Publicado: (2025)
por: Lu, Ye-Xin, et al.
Publicado: (2025)
A Comprehensive Study of the Current State-of-the-Art in Nepali Automatic Speech Recognition Systems
por: Ghimire, Rupak Raj, et al.
Publicado: (2024)
por: Ghimire, Rupak Raj, et al.
Publicado: (2024)
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception
por: Zhang, Jiawei, et al.
Publicado: (2024)
por: Zhang, Jiawei, et al.
Publicado: (2024)
Exploring synthetic data for cross-speaker style transfer in style representation based TTS
por: Ueda, Lucas H., et al.
Publicado: (2024)
por: Ueda, Lucas H., et al.
Publicado: (2024)
Ejemplares similares
-
Exploring an Inter-Pausal Unit (IPU) based Approach for Indic End-to-End TTS Systems
por: Prakash, Anusha, et al.
Publicado: (2024) -
AutoProsody: A Prosodic Feature Extraction Tool for Indian Languages
por: Thinakaran, Preethi, et al.
Publicado: (2025) -
A Unified Framework for Collecting Text-to-Speech Synthesis Datasets for 22 Indian Languages
por: Sathiyamoorthy, Sujitha, et al.
Publicado: (2024) -
Intelli-Z: Toward Intelligible Zero-Shot TTS
por: Jung, Sunghee, et al.
Publicado: (2024) -
SponTTS: modeling and transferring spontaneous style for TTS
por: Li, Hanzhao, et al.
Publicado: (2023)