Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xinsheng, Jiang, Mingqi, Ma, Ziyang, Zhang, Ziyu, Liu, Songxiang, Li, Linqin, Liang, Zheng, Zheng, Qixi, Wang, Rui, Feng, Xiaoqin, Bian, Weizhen, Ye, Zhen, Cheng, Sitong, Yuan, Ruibin, Zhao, Zhixian, Zhu, Xinfa, Pan, Jiahao, Xue, Liumeng, Zhu, Pengcheng, Chen, Yunlin, Li, Zhifei, Chen, Xie, Xie, Lei, Guo, Yike, Xue, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
von: Li, Hanzhao, et al.
Veröffentlicht: (2024)
von: Li, Hanzhao, et al.
Veröffentlicht: (2024)
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
Text-aware and Context-aware Expressive Audiobook Speech Synthesis
von: Guo, Dake, et al.
Veröffentlicht: (2024)
von: Guo, Dake, et al.
Veröffentlicht: (2024)
Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
von: Ye, Zhen, et al.
Veröffentlicht: (2025)
von: Ye, Zhen, et al.
Veröffentlicht: (2025)
Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought
von: Zhao, Zhixian, et al.
Veröffentlicht: (2025)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2025)
WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
Accent-VITS:accent transfer for end-to-end TTS
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
von: Xia, Kangxiang, et al.
Veröffentlicht: (2025)
von: Xia, Kangxiang, et al.
Veröffentlicht: (2025)
Transfer the linguistic representations from TTS to accent conversion with non-parallel data
von: Chen, Xi, et al.
Veröffentlicht: (2024)
von: Chen, Xi, et al.
Veröffentlicht: (2024)
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
von: Zheng, Qixi, et al.
Veröffentlicht: (2025)
von: Zheng, Qixi, et al.
Veröffentlicht: (2025)
OpenSTBench: Beyond Semantic Evaluation for Speech Translation
von: An, Yanjie, et al.
Veröffentlicht: (2026)
von: An, Yanjie, et al.
Veröffentlicht: (2026)
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement
von: Kang, Boyi, et al.
Veröffentlicht: (2025)
von: Kang, Boyi, et al.
Veröffentlicht: (2025)
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis
von: Liu, Qingyu, et al.
Veröffentlicht: (2025)
von: Liu, Qingyu, et al.
Veröffentlicht: (2025)
NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations
von: Xue, Liumeng, et al.
Veröffentlicht: (2026)
von: Xue, Liumeng, et al.
Veröffentlicht: (2026)
ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing
von: Chen, Xi, et al.
Veröffentlicht: (2026)
von: Chen, Xi, et al.
Veröffentlicht: (2026)
KALL-E:Autoregressive Speech Synthesis with Next-Distribution Prediction
von: Xia, Kangxiang, et al.
Veröffentlicht: (2024)
von: Xia, Kangxiang, et al.
Veröffentlicht: (2024)
AudioX: A Unified Framework for Anything-to-Audio Generation
von: Tian, Zeyue, et al.
Veröffentlicht: (2025)
von: Tian, Zeyue, et al.
Veröffentlicht: (2025)
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
FlashSpeech: Efficient Zero-Shot Speech Synthesis
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
von: Li, Hanzhao, et al.
Veröffentlicht: (2025)
von: Li, Hanzhao, et al.
Veröffentlicht: (2025)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
SELM: Speech Enhancement Using Discrete Tokens and Language Models
von: Wang, Ziqian, et al.
Veröffentlicht: (2023)
von: Wang, Ziqian, et al.
Veröffentlicht: (2023)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2025)
Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities
von: Xie, Zhifei, et al.
Veröffentlicht: (2024)
von: Xie, Zhifei, et al.
Veröffentlicht: (2024)
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
von: Chen, Huakang, et al.
Veröffentlicht: (2026)
von: Chen, Huakang, et al.
Veröffentlicht: (2026)
WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation
von: Li, Longhao, et al.
Veröffentlicht: (2025)
von: Li, Longhao, et al.
Veröffentlicht: (2025)
StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations
von: Liu, Sen, et al.
Veröffentlicht: (2024)
von: Liu, Sen, et al.
Veröffentlicht: (2024)
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
von: Liu, Huadai, et al.
Veröffentlicht: (2023)
Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis
von: Feng, Pengchao, et al.
Veröffentlicht: (2025)
von: Feng, Pengchao, et al.
Veröffentlicht: (2025)
LLaDA-TTS: Unifying Speech Synthesis and Zero-Shot Editing via Masked Diffusion Modeling
von: Fan, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Fan, Xiaoyu, et al.
Veröffentlicht: (2026)
MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
von: Li, Haitao, et al.
Veröffentlicht: (2026)
von: Li, Haitao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023) -
Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
von: Li, Hanzhao, et al.
Veröffentlicht: (2024) -
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
von: Cheng, Sitong, et al.
Veröffentlicht: (2025) -
Text-aware and Context-aware Expressive Audiobook Speech Synthesis
von: Guo, Dake, et al.
Veröffentlicht: (2024) -
Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
von: Ye, Zhen, et al.
Veröffentlicht: (2025)