Next Tokens Denoising for Speech Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Yanqing, Xue, Ruiqing, Zhang, Chong, Liu, Yufei, Wang, Gang, Li, Bohan, Qian, Yao, He, Lei, Liu, Shujie, Zhao, Sheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
von: Yuan, Ze, et al.
Veröffentlicht: (2024)
von: Yuan, Ze, et al.
Veröffentlicht: (2024)
Autoregressive Speech Synthesis without Vector Quantization
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
von: Han, Bing, et al.
Veröffentlicht: (2024)
von: Han, Bing, et al.
Veröffentlicht: (2024)
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
von: Chen, Sanyuan, et al.
Veröffentlicht: (2024)
von: Chen, Sanyuan, et al.
Veröffentlicht: (2024)
A Unified Neural Codec Language Model for Selective Editable Text to Speech Generation
von: Pei, Hanchen, et al.
Veröffentlicht: (2026)
von: Pei, Hanchen, et al.
Veröffentlicht: (2026)
Boosting Large Language Model for Speech Synthesis: An Empirical Study
von: Hao, Hongkun, et al.
Veröffentlicht: (2023)
von: Hao, Hongkun, et al.
Veröffentlicht: (2023)
KALL-E:Autoregressive Speech Synthesis with Next-Distribution Prediction
von: Xia, Kangxiang, et al.
Veröffentlicht: (2024)
von: Xia, Kangxiang, et al.
Veröffentlicht: (2024)
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
von: Le, Chenyang, et al.
Veröffentlicht: (2024)
von: Le, Chenyang, et al.
Veröffentlicht: (2024)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
von: Wang, Huimeng, et al.
Veröffentlicht: (2025)
von: Wang, Huimeng, et al.
Veröffentlicht: (2025)
Closing the Modality Reasoning Gap for Speech Large Language Models
von: Wang, Chaoren, et al.
Veröffentlicht: (2026)
von: Wang, Chaoren, et al.
Veröffentlicht: (2026)
Advanced Long-Content Speech Recognition With Factorized Neural Transducer
von: Gong, Xun, et al.
Veröffentlicht: (2024)
von: Gong, Xun, et al.
Veröffentlicht: (2024)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
von: Cui, Yang, et al.
Veröffentlicht: (2025)
von: Cui, Yang, et al.
Veröffentlicht: (2025)
Continuous Speech Tokenizer in Text To Speech
von: Li, Yixing, et al.
Veröffentlicht: (2024)
von: Li, Yixing, et al.
Veröffentlicht: (2024)
Generative Expressive Conversational Speech Synthesis
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Pairwise Evaluation of Accent Similarity in Speech Synthesis
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2025)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2025)
Intra- and Inter-modal Context Interaction Modeling for Conversational Speech Synthesis
von: Jia, Zhenqi, et al.
Veröffentlicht: (2024)
von: Jia, Zhenqi, et al.
Veröffentlicht: (2024)
Investigating Group Relative Policy Optimization for Diffusion Transformer based Text-to-Audio Generation
von: Gu, Yi, et al.
Veröffentlicht: (2026)
von: Gu, Yi, et al.
Veröffentlicht: (2026)
SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech Recognition
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
von: Xue, Hongfei, et al.
Veröffentlicht: (2023)
Frontend Token Enhancement for Token-Based Speech Recognition
von: Ashihara, Takanori, et al.
Veröffentlicht: (2026)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2026)
Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
von: Li, Bohan, et al.
Veröffentlicht: (2025)
von: Li, Bohan, et al.
Veröffentlicht: (2025)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
Speech Denoising with Auditory Models
von: Saddler, Mark R., et al.
Veröffentlicht: (2020)
von: Saddler, Mark R., et al.
Veröffentlicht: (2020)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
STAB: Speech Tokenizer Assessment Benchmark
von: Vashishth, Shikhar, et al.
Veröffentlicht: (2024)
von: Vashishth, Shikhar, et al.
Veröffentlicht: (2024)
Exploring the Potential of Large Multimodal Models as Effective Alternatives for Pronunciation Assessment
von: Wang, Ke, et al.
Veröffentlicht: (2025)
von: Wang, Ke, et al.
Veröffentlicht: (2025)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
Speech Editing -- a Summary
von: Kässmann, Tobias, et al.
Veröffentlicht: (2024)
von: Kässmann, Tobias, et al.
Veröffentlicht: (2024)
StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
von: Guo, Dake, et al.
Veröffentlicht: (2025)
von: Guo, Dake, et al.
Veröffentlicht: (2025)
CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker Conversations
von: Zhang, Leying, et al.
Veröffentlicht: (2024)
von: Zhang, Leying, et al.
Veröffentlicht: (2024)
RaD-Net: A Repairing and Denoising Network for Speech Signal Improvement
von: Liu, Mingshuai, et al.
Veröffentlicht: (2024)
von: Liu, Mingshuai, et al.
Veröffentlicht: (2024)
WavMark: Watermarking for Audio Generation
von: Chen, Guangyu, et al.
Veröffentlicht: (2023)
von: Chen, Guangyu, et al.
Veröffentlicht: (2023)
SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
Factorized RVQ-GAN For Disentangled Speech Tokenization
von: Khurana, Sameer, et al.
Veröffentlicht: (2025)
von: Khurana, Sameer, et al.
Veröffentlicht: (2025)
Benchmarking Prosody Encoding in Discrete Speech Tokens
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
LAST: Language Model Aware Speech Tokenization
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
WavLLM: Towards Robust and Adaptive Speech Large Language Model
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
von: Yuan, Ze, et al.
Veröffentlicht: (2024) -
Autoregressive Speech Synthesis without Vector Quantization
von: Meng, Lingwei, et al.
Veröffentlicht: (2024) -
FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
von: Wang, Hui, et al.
Veröffentlicht: (2025) -
VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
von: Han, Bing, et al.
Veröffentlicht: (2024) -
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
von: Chen, Sanyuan, et al.
Veröffentlicht: (2024)