MahaTTS: A Unified Framework for Multilingual Text-to-Speech Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Singh, Jaskaran, Chowdhury, Amartya Roy, Prabhakar, Raghav, W, Varshul C. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MunTTS: A Text-to-Speech System for Mundari
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers and Consistency Models
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations
von: Liu, Sen, et al.
Veröffentlicht: (2024)
von: Liu, Sen, et al.
Veröffentlicht: (2024)
GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
von: Lee, Seokgi, et al.
Veröffentlicht: (2025)
von: Lee, Seokgi, et al.
Veröffentlicht: (2025)
Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation
von: Di, Xinhan, et al.
Veröffentlicht: (2024)
von: Di, Xinhan, et al.
Veröffentlicht: (2024)
DiaMoE-TTS: A Unified IPA-Based Dialect TTS Framework with Mixture-of-Experts and Parameter-Efficient Zero-Shot Adaptation
von: Chen, Ziqi, et al.
Veröffentlicht: (2025)
von: Chen, Ziqi, et al.
Veröffentlicht: (2025)
TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis
von: Wang, Xi, et al.
Veröffentlicht: (2026)
von: Wang, Xi, et al.
Veröffentlicht: (2026)
TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward
von: Wang, Guansu, et al.
Veröffentlicht: (2025)
von: Wang, Guansu, et al.
Veröffentlicht: (2025)
Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation
von: Lou, Haowei, et al.
Veröffentlicht: (2025)
von: Lou, Haowei, et al.
Veröffentlicht: (2025)
XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
HyperTTS: Parameter Efficient Adaptation in Text to Speech using Hypernetworks
von: Li, Yingting, et al.
Veröffentlicht: (2024)
von: Li, Yingting, et al.
Veröffentlicht: (2024)
Indonesian-English Code-Switching Speech Synthesizer Utilizing Multilingual STEN-TTS and Bert LID
von: Handoyo, Ahmad Alfani, et al.
Veröffentlicht: (2024)
von: Handoyo, Ahmad Alfani, et al.
Veröffentlicht: (2024)
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
von: He, Haorui, et al.
Veröffentlicht: (2024)
von: He, Haorui, et al.
Veröffentlicht: (2024)
Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness
von: Feng, Xincan, et al.
Veröffentlicht: (2024)
von: Feng, Xincan, et al.
Veröffentlicht: (2024)
GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM
von: Song, Yaodong, et al.
Veröffentlicht: (2025)
von: Song, Yaodong, et al.
Veröffentlicht: (2025)
FMSD-TTS: Few-shot Multi-Speaker Multi-Dialect Text-to-Speech Synthesis for Ü-Tsang, Amdo and Kham Speech Dataset Generation
von: Liu, Yutong, et al.
Veröffentlicht: (2025)
von: Liu, Yutong, et al.
Veröffentlicht: (2025)
SpeechTaxi: On Multilingual Semantic Speech Classification
von: Keller, Lennart, et al.
Veröffentlicht: (2024)
von: Keller, Lennart, et al.
Veröffentlicht: (2024)
A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations
von: Saengthong, Phurich, et al.
Veröffentlicht: (2025)
von: Saengthong, Phurich, et al.
Veröffentlicht: (2025)
Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond
von: Lee, Beomseok, et al.
Veröffentlicht: (2024)
von: Lee, Beomseok, et al.
Veröffentlicht: (2024)
Enhancing Multilingual Voice Toxicity Detection with Speech-Text Alignment
von: Liu, Joseph, et al.
Veröffentlicht: (2024)
von: Liu, Joseph, et al.
Veröffentlicht: (2024)
EE-TTS: Emphatic Expressive TTS with Linguistic Information
von: Zhong, Yi, et al.
Veröffentlicht: (2023)
von: Zhong, Yi, et al.
Veröffentlicht: (2023)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis
von: Abilbekov, Adal, et al.
Veröffentlicht: (2024)
von: Abilbekov, Adal, et al.
Veröffentlicht: (2024)
Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
von: Yang, Yifan, et al.
Veröffentlicht: (2025)
von: Yang, Yifan, et al.
Veröffentlicht: (2025)
Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems
von: Allbert, Rumi, et al.
Veröffentlicht: (2025)
von: Allbert, Rumi, et al.
Veröffentlicht: (2025)
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
UtterTune: LoRA-Based Target-Language Pronunciation Edit and Control in Multilingual Text-to-Speech
von: Kato, Shuhei
Veröffentlicht: (2025)
von: Kato, Shuhei
Veröffentlicht: (2025)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
von: Łajszczak, Mateusz, et al.
Veröffentlicht: (2024)
von: Łajszczak, Mateusz, et al.
Veröffentlicht: (2024)
A2TTS: TTS for Low Resource Indian Languages
von: Bhadoriya, Ayush Singh, et al.
Veröffentlicht: (2025)
von: Bhadoriya, Ayush Singh, et al.
Veröffentlicht: (2025)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2024)
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2024)
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
von: Li, Fengjin, et al.
Veröffentlicht: (2025)
von: Li, Fengjin, et al.
Veröffentlicht: (2025)
VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Textless Unit-to-Unit training for Many-to-Many Multilingual Speech-to-Speech Translation
von: Kim, Minsu, et al.
Veröffentlicht: (2023)
von: Kim, Minsu, et al.
Veröffentlicht: (2023)
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
von: Li, Zhu, et al.
Veröffentlicht: (2025)
von: Li, Zhu, et al.
Veröffentlicht: (2025)
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MunTTS: A Text-to-Speech System for Mundari
von: Gumma, Varun, et al.
Veröffentlicht: (2024) -
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025) -
CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers and Consistency Models
von: Li, Xiang, et al.
Veröffentlicht: (2024) -
StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations
von: Liu, Sen, et al.
Veröffentlicht: (2024) -
GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
von: Lee, Seokgi, et al.
Veröffentlicht: (2025)