TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Xingchen, Xing, Mengtao, Ma, Changwei, Li, Shengqiang, Wu, Di, Zhang, Binbin, Pan, Fuping, Zhou, Dinghao, Zhang, Yuekai, Lei, Shun, Peng, Zhendong, Wu, Zhiyong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
von: Song, Xingchen, et al.
Veröffentlicht: (2024)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024)
E1 TTS: Simple and Fast Non-Autoregressive TTS
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
CAST-TTS: A Simple Cross-Attention Framework for Unified Timbre Control in TTS
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024
von: Zhou, Shuoyi, et al.
Veröffentlicht: (2024)
von: Zhou, Shuoyi, et al.
Veröffentlicht: (2024)
Bridging the gap between training and inference in LM-based TTS models
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information
von: Wang, Rui, et al.
Veröffentlicht: (2025)
von: Wang, Rui, et al.
Veröffentlicht: (2025)
EE-TTS: Emphatic Expressive TTS with Linguistic Information
von: Zhong, Yi, et al.
Veröffentlicht: (2023)
von: Zhong, Yi, et al.
Veröffentlicht: (2023)
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
von: Qharabagh, Mahta Fetrat, et al.
Veröffentlicht: (2024)
von: Qharabagh, Mahta Fetrat, et al.
Veröffentlicht: (2024)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
von: Xue, Heyang, et al.
Veröffentlicht: (2025)
von: Xue, Heyang, et al.
Veröffentlicht: (2025)
Accent-VITS:accent transfer for end-to-end TTS
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS
von: Nie, Sihang, et al.
Veröffentlicht: (2025)
von: Nie, Sihang, et al.
Veröffentlicht: (2025)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis
von: Chen, Zhiyong, et al.
Veröffentlicht: (2024)
von: Chen, Zhiyong, et al.
Veröffentlicht: (2024)
SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word Level
von: Tee, Hitomi Jin Ling, et al.
Veröffentlicht: (2025)
von: Tee, Hitomi Jin Ling, et al.
Veröffentlicht: (2025)
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
Qwen3-TTS Technical Report
von: Hu, Hangrui, et al.
Veröffentlicht: (2026)
von: Hu, Hangrui, et al.
Veröffentlicht: (2026)
A Dataset for Automatic Assessment of TTS Quality in Spanish
von: Welford, Alejandro Sosa, et al.
Veröffentlicht: (2025)
von: Welford, Alejandro Sosa, et al.
Veröffentlicht: (2025)
Intelli-Z: Toward Intelligible Zero-Shot TTS
von: Jung, Sunghee, et al.
Veröffentlicht: (2024)
von: Jung, Sunghee, et al.
Veröffentlicht: (2024)
Zero-shot Cross-lingual Voice Transfer for TTS
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
von: Biadsy, Fadi, et al.
Veröffentlicht: (2024)
DQR-TTS: Semi-supervised Text-to-speech Synthesis with Dynamic Quantized Representation
von: Wang, Jianzong, et al.
Veröffentlicht: (2023)
von: Wang, Jianzong, et al.
Veröffentlicht: (2023)
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2024)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2025)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2025)
Enhancing TTS Stability in Hebrew using Discrete Semantic Units
von: Zeldes, Ella, et al.
Veröffentlicht: (2024)
von: Zeldes, Ella, et al.
Veröffentlicht: (2024)
Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
von: He, Xinlu, et al.
Veröffentlicht: (2025)
von: He, Xinlu, et al.
Veröffentlicht: (2025)
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
SPAM: Style Prompt Adherence Metric for Prompt-based TTS
von: Cho, Chanhee, et al.
Veröffentlicht: (2026)
von: Cho, Chanhee, et al.
Veröffentlicht: (2026)
Low-Resource Self-Supervised Learning with SSL-Enhanced TTS
von: Hsu, Po-chun, et al.
Veröffentlicht: (2023)
von: Hsu, Po-chun, et al.
Veröffentlicht: (2023)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
Differentiable Reward Optimization for LLM based TTS system
von: Gao, Changfeng, et al.
Veröffentlicht: (2025)
von: Gao, Changfeng, et al.
Veröffentlicht: (2025)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
von: Ren, Yong, et al.
Veröffentlicht: (2026)
von: Ren, Yong, et al.
Veröffentlicht: (2026)
Prosodic Parameter Manipulation in TTS generated speech for Controlled Speech Generation
von: Chary, Podakanti Satyajith
Veröffentlicht: (2024)
von: Chary, Podakanti Satyajith
Veröffentlicht: (2024)
F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization
von: Sun, Xiaohui, et al.
Veröffentlicht: (2025)
von: Sun, Xiaohui, et al.
Veröffentlicht: (2025)
Scaling NVIDIA's Multi-speaker Multi-lingual TTS Systems with Zero-Shot TTS to Indic Languages
von: Arora, Akshit, et al.
Veröffentlicht: (2024)
von: Arora, Akshit, et al.
Veröffentlicht: (2024)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch
von: Song, Xingchen, et al.
Veröffentlicht: (2024) -
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
von: Eskimez, Sefik Emre, et al.
Veröffentlicht: (2024) -
E1 TTS: Simple and Fast Non-Autoregressive TTS
von: Liu, Zhijun, et al.
Veröffentlicht: (2024) -
CAST-TTS: A Simple Cross-Attention Framework for Unified Timbre Control in TTS
von: Zheng, Zihao, et al.
Veröffentlicht: (2026) -
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)