Bridging the gap between training and inference in LM-based TTS models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Ruonan, Mu, Lingzhou, Wu, Xixin, Zhang, Kai |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
di: Wu, Wenxuan, et al.
Pubblicazione: (2024)
di: Wu, Wenxuan, et al.
Pubblicazione: (2024)
SponTTS: modeling and transferring spontaneous style for TTS
di: Li, Hanzhao, et al.
Pubblicazione: (2023)
di: Li, Hanzhao, et al.
Pubblicazione: (2023)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets
di: Liu, Chenlin, et al.
Pubblicazione: (2025)
di: Liu, Chenlin, et al.
Pubblicazione: (2025)
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
di: Xue, Heyang, et al.
Pubblicazione: (2025)
di: Xue, Heyang, et al.
Pubblicazione: (2025)
E1 TTS: Simple and Fast Non-Autoregressive TTS
di: Liu, Zhijun, et al.
Pubblicazione: (2024)
di: Liu, Zhijun, et al.
Pubblicazione: (2024)
MBCodec:Thorough disentangle for high-fidelity audio compression
di: Zhang, Ruonan, et al.
Pubblicazione: (2025)
di: Zhang, Ruonan, et al.
Pubblicazione: (2025)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
di: Chen, Weidong, et al.
Pubblicazione: (2025)
di: Chen, Weidong, et al.
Pubblicazione: (2025)
SPAM: Style Prompt Adherence Metric for Prompt-based TTS
di: Cho, Chanhee, et al.
Pubblicazione: (2026)
di: Cho, Chanhee, et al.
Pubblicazione: (2026)
Accent-VITS:accent transfer for end-to-end TTS
di: Ma, Linhan, et al.
Pubblicazione: (2023)
di: Ma, Linhan, et al.
Pubblicazione: (2023)
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
di: Eskimez, Sefik Emre, et al.
Pubblicazione: (2024)
di: Eskimez, Sefik Emre, et al.
Pubblicazione: (2024)
ManaTTS Persian: a recipe for creating TTS datasets for lower resource languages
di: Qharabagh, Mahta Fetrat, et al.
Pubblicazione: (2024)
di: Qharabagh, Mahta Fetrat, et al.
Pubblicazione: (2024)
AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
di: Wang, Yuanyuan, et al.
Pubblicazione: (2024)
di: Wang, Yuanyuan, et al.
Pubblicazione: (2024)
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception
di: Zhang, Jiawei, et al.
Pubblicazione: (2024)
di: Zhang, Jiawei, et al.
Pubblicazione: (2024)
TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
di: Song, Xingchen, et al.
Pubblicazione: (2024)
di: Song, Xingchen, et al.
Pubblicazione: (2024)
The Codec Language Model-based Zero-Shot Spontaneous Style TTS System for CoVoC Challenge 2024
di: Zhou, Shuoyi, et al.
Pubblicazione: (2024)
di: Zhou, Shuoyi, et al.
Pubblicazione: (2024)
F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization
di: Sun, Xiaohui, et al.
Pubblicazione: (2025)
di: Sun, Xiaohui, et al.
Pubblicazione: (2025)
DQR-TTS: Semi-supervised Text-to-speech Synthesis with Dynamic Quantized Representation
di: Wang, Jianzong, et al.
Pubblicazione: (2023)
di: Wang, Jianzong, et al.
Pubblicazione: (2023)
Adversarial training of Keyword Spotting to Minimize TTS Data Overfitting
di: Park, Hyun Jin, et al.
Pubblicazione: (2024)
di: Park, Hyun Jin, et al.
Pubblicazione: (2024)
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
di: Anastassiou, Philip, et al.
Pubblicazione: (2024)
di: Anastassiou, Philip, et al.
Pubblicazione: (2024)
Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
di: Guo, Haohan, et al.
Pubblicazione: (2024)
di: Guo, Haohan, et al.
Pubblicazione: (2024)
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
di: Chen, Yushen, et al.
Pubblicazione: (2024)
di: Chen, Yushen, et al.
Pubblicazione: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
di: Guo, Yinlin, et al.
Pubblicazione: (2024)
di: Guo, Yinlin, et al.
Pubblicazione: (2024)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
di: Guo, Hao-Han, et al.
Pubblicazione: (2024)
di: Guo, Hao-Han, et al.
Pubblicazione: (2024)
HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS
di: Nie, Sihang, et al.
Pubblicazione: (2025)
di: Nie, Sihang, et al.
Pubblicazione: (2025)
Exploring synthetic data for cross-speaker style transfer in style representation based TTS
di: Ueda, Lucas H., et al.
Pubblicazione: (2024)
di: Ueda, Lucas H., et al.
Pubblicazione: (2024)
Are Transformers in Pre-trained LM A Good ASR Encoder? An Empirical Study
di: An, Keyu, et al.
Pubblicazione: (2024)
di: An, Keyu, et al.
Pubblicazione: (2024)
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information
di: Wang, Rui, et al.
Pubblicazione: (2025)
di: Wang, Rui, et al.
Pubblicazione: (2025)
Differentiable Reward Optimization for LLM based TTS system
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
di: Gao, Changfeng, et al.
Pubblicazione: (2025)
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
di: Guan, Wenhao, et al.
Pubblicazione: (2025)
di: Guan, Wenhao, et al.
Pubblicazione: (2025)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
di: Guan, Wenhao, et al.
Pubblicazione: (2023)
di: Guan, Wenhao, et al.
Pubblicazione: (2023)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
di: Jiang, Ziyue, et al.
Pubblicazione: (2023)
di: Jiang, Ziyue, et al.
Pubblicazione: (2023)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
di: Guo, Hao-Han, et al.
Pubblicazione: (2025)
di: Guo, Hao-Han, et al.
Pubblicazione: (2025)
EE-TTS: Emphatic Expressive TTS with Linguistic Information
di: Zhong, Yi, et al.
Pubblicazione: (2023)
di: Zhong, Yi, et al.
Pubblicazione: (2023)
A Dataset for Automatic Assessment of TTS Quality in Spanish
di: Welford, Alejandro Sosa, et al.
Pubblicazione: (2025)
di: Welford, Alejandro Sosa, et al.
Pubblicazione: (2025)
Intelli-Z: Toward Intelligible Zero-Shot TTS
di: Jung, Sunghee, et al.
Pubblicazione: (2024)
di: Jung, Sunghee, et al.
Pubblicazione: (2024)
Zero-shot Cross-lingual Voice Transfer for TTS
di: Biadsy, Fadi, et al.
Pubblicazione: (2024)
di: Biadsy, Fadi, et al.
Pubblicazione: (2024)
EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge
di: Manku, Ruskin Raj, et al.
Pubblicazione: (2025)
di: Manku, Ruskin Raj, et al.
Pubblicazione: (2025)
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
di: Lu, Ye-Xin, et al.
Pubblicazione: (2025)
di: Lu, Ye-Xin, et al.
Pubblicazione: (2025)
STSM-FiLM: A FiLM-Conditioned Neural Architecture for Time-Scale Modification of Speech
di: Wisnu, Dyah A. M. G., et al.
Pubblicazione: (2025)
di: Wisnu, Dyah A. M. G., et al.
Pubblicazione: (2025)
Documenti analoghi
-
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
di: Wu, Wenxuan, et al.
Pubblicazione: (2024) -
SponTTS: modeling and transferring spontaneous style for TTS
di: Li, Hanzhao, et al.
Pubblicazione: (2023) -
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
di: Wang, Yuanyuan, et al.
Pubblicazione: (2025) -
Mitigating Hallucinations in LM-Based TTS Models via Distribution Alignment Using GFlowNets
di: Liu, Chenlin, et al.
Pubblicazione: (2025) -
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
di: Xue, Heyang, et al.
Pubblicazione: (2025)