Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Anand, Srija, Varadhan, Praveen Srinivasa, Sankar, Ashwin, Raju, Giri, Khapra, Mitesh M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2024)
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2024)
Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
ELAICHI: Enhancing Low-resource TTS by Addressing Infrequent and Low-frequency Character Bigrams
von: Anand, Srija, et al.
Veröffentlicht: (2024)
von: Anand, Srija, et al.
Veröffentlicht: (2024)
The State Of TTS: A Case Study with Human Fooling Rates
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2024)
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2024)
IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS
von: Sankar, Ashwin, et al.
Veröffentlicht: (2024)
von: Sankar, Ashwin, et al.
Veröffentlicht: (2024)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
von: Ren, Yong, et al.
Veröffentlicht: (2026)
von: Ren, Yong, et al.
Veröffentlicht: (2026)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
von: Xue, Heyang, et al.
Veröffentlicht: (2025)
von: Xue, Heyang, et al.
Veröffentlicht: (2025)
A2TTS: TTS for Low Resource Indian Languages
von: Bhadoriya, Ayush Singh, et al.
Veröffentlicht: (2025)
von: Bhadoriya, Ayush Singh, et al.
Veröffentlicht: (2025)
Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
von: He, Xinlu, et al.
Veröffentlicht: (2025)
von: He, Xinlu, et al.
Veröffentlicht: (2025)
Empowering Low-Resource Language ASR via Large-Scale Pseudo Labeling
von: Bhogale, Kaushal Santosh, et al.
Veröffentlicht: (2024)
von: Bhogale, Kaushal Santosh, et al.
Veröffentlicht: (2024)
Low-Resource Self-Supervised Learning with SSL-Enhanced TTS
von: Hsu, Po-chun, et al.
Veröffentlicht: (2023)
von: Hsu, Po-chun, et al.
Veröffentlicht: (2023)
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
E1 TTS: Simple and Fast Non-Autoregressive TTS
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
Towards Developing State-of-the-Art TTS Synthesisers for 13 Indian Languages with Signal Processing aided Alignments
von: Prakash, Anusha, et al.
Veröffentlicht: (2022)
von: Prakash, Anusha, et al.
Veröffentlicht: (2022)
Empowering Communication: Speech Technology for Indian and Western Accents through AI-powered Speech Synthesis
von: R, Vinotha, et al.
Veröffentlicht: (2024)
von: R, Vinotha, et al.
Veröffentlicht: (2024)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
von: Qharabagh, Mahta Fetrat, et al.
Veröffentlicht: (2024)
von: Qharabagh, Mahta Fetrat, et al.
Veröffentlicht: (2024)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
Accent-VITS:accent transfer for end-to-end TTS
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
A Dataset for Automatic Assessment of TTS Quality in Spanish
von: Welford, Alejandro Sosa, et al.
Veröffentlicht: (2025)
von: Welford, Alejandro Sosa, et al.
Veröffentlicht: (2025)
Intelli-Z: Toward Intelligible Zero-Shot TTS
von: Jung, Sunghee, et al.
Veröffentlicht: (2024)
von: Jung, Sunghee, et al.
Veröffentlicht: (2024)
Zero-shot Cross-lingual Voice Transfer for TTS
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
A Neural Score Follower for Computer Accompaniment of Polyphonic Musical Instruments
von: Pillay, Ashwin
Veröffentlicht: (2025)
von: Pillay, Ashwin
Veröffentlicht: (2025)
SelfTTS: cross-speaker style transfer through explicit embedding disentanglement and self-refinement using self-augmentation
von: Ueda, Lucas H., et al.
Veröffentlicht: (2026)
von: Ueda, Lucas H., et al.
Veröffentlicht: (2026)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2025)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2025)
Enhancing TTS Stability in Hebrew using Discrete Semantic Units
von: Zeldes, Ella, et al.
Veröffentlicht: (2024)
von: Zeldes, Ella, et al.
Veröffentlicht: (2024)
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
SPAM: Style Prompt Adherence Metric for Prompt-based TTS
von: Cho, Chanhee, et al.
Veröffentlicht: (2026)
von: Cho, Chanhee, et al.
Veröffentlicht: (2026)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
Bridging the gap between training and inference in LM-based TTS models
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
Prosodic Parameter Manipulation in TTS generated speech for Controlled Speech Generation
von: Chary, Podakanti Satyajith
Veröffentlicht: (2024)
von: Chary, Podakanti Satyajith
Veröffentlicht: (2024)
LoRP-TTS: Low-Rank Personalized Text-To-Speech
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2025)
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2025)
EE-TTS: Emphatic Expressive TTS with Linguistic Information
von: Zhong, Yi, et al.
Veröffentlicht: (2023)
von: Zhong, Yi, et al.
Veröffentlicht: (2023)
Transmission of High-Amplitude Sound through Leakages of Ill-fitting Earplugs
von: Yu, Haocheng, et al.
Veröffentlicht: (2025)
von: Yu, Haocheng, et al.
Veröffentlicht: (2025)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
von: Peng, Puyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2024) -
Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025) -
ELAICHI: Enhancing Low-resource TTS by Addressing Infrequent and Low-frequency Character Bigrams
von: Anand, Srija, et al.
Veröffentlicht: (2024) -
The State Of TTS: A Case Study with Human Fooling Rates
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025) -
Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2024)