Decoding Order Matters in Autoregressive Speech Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhao, Minghui, Ragni, Anton |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
di: Lin, Zijian, et al.
Pubblicazione: (2025)
di: Lin, Zijian, et al.
Pubblicazione: (2025)
Unifying Speech Recognition, Synthesis and Conversion with Autoregressive Transformers
di: Cai, Runyuan, et al.
Pubblicazione: (2026)
di: Cai, Runyuan, et al.
Pubblicazione: (2026)
Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users using Intermediate ASR Features and Human Memory Models
di: Mogridge, Rhiannon, et al.
Pubblicazione: (2024)
di: Mogridge, Rhiannon, et al.
Pubblicazione: (2024)
Beyond the Utterance: An Empirical Study of Very Long Context Speech Recognition
di: Flynn, Robert, et al.
Pubblicazione: (2026)
di: Flynn, Robert, et al.
Pubblicazione: (2026)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
di: Leung, Wing-Zin, et al.
Pubblicazione: (2024)
di: Leung, Wing-Zin, et al.
Pubblicazione: (2024)
Affect Decoding in Phonated and Silent Speech Production from Surface EMG
di: Pistrosch, Simon, et al.
Pubblicazione: (2026)
di: Pistrosch, Simon, et al.
Pubblicazione: (2026)
Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
di: Nguyen, Tan Dat, et al.
Pubblicazione: (2024)
di: Nguyen, Tan Dat, et al.
Pubblicazione: (2024)
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
di: Lee, Seo-Hyun, et al.
Pubblicazione: (2023)
Towards Unified Neural Decoding of Perceived, Spoken and Imagined Speech from EEG Signals
di: Lee, Jung-Sun, et al.
Pubblicazione: (2024)
di: Lee, Jung-Sun, et al.
Pubblicazione: (2024)
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
di: Moell, Birger, et al.
Pubblicazione: (2025)
di: Moell, Birger, et al.
Pubblicazione: (2025)
Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models
di: Shah, Neil, et al.
Pubblicazione: (2024)
di: Shah, Neil, et al.
Pubblicazione: (2024)
MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI
di: Shah, Neil, et al.
Pubblicazione: (2024)
di: Shah, Neil, et al.
Pubblicazione: (2024)
Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding
di: Ahn, Hoseong, et al.
Pubblicazione: (2026)
di: Ahn, Hoseong, et al.
Pubblicazione: (2026)
Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis
di: Juvela, Lauri, et al.
Pubblicazione: (2024)
di: Juvela, Lauri, et al.
Pubblicazione: (2024)
AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis
di: Cao, Yubing, et al.
Pubblicazione: (2025)
di: Cao, Yubing, et al.
Pubblicazione: (2025)
Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
di: Yang, Dong, et al.
Pubblicazione: (2025)
di: Yang, Dong, et al.
Pubblicazione: (2025)
Lina-Speech: Gated Linear Attention and Initial-State Tuning for Multi-Sample Prompting Text-To-Speech Synthesis
di: Lemerle, Théodor, et al.
Pubblicazione: (2024)
di: Lemerle, Théodor, et al.
Pubblicazione: (2024)
Parallel Synthesis for Autoregressive Speech Generation
di: Hsu, Po-chun, et al.
Pubblicazione: (2022)
di: Hsu, Po-chun, et al.
Pubblicazione: (2022)
A Mel Spectrogram Enhancement Paradigm Based on CWT in Speech Synthesis
di: Hu, Guoqiang, et al.
Pubblicazione: (2024)
di: Hu, Guoqiang, et al.
Pubblicazione: (2024)
Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment
di: Neekhara, Paarth, et al.
Pubblicazione: (2024)
di: Neekhara, Paarth, et al.
Pubblicazione: (2024)
DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models
di: wu, Weihao, et al.
Pubblicazione: (2025)
di: wu, Weihao, et al.
Pubblicazione: (2025)
ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning
di: Zhu, Tao, et al.
Pubblicazione: (2025)
di: Zhu, Tao, et al.
Pubblicazione: (2025)
Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer
di: Wang, Yongqi, et al.
Pubblicazione: (2023)
di: Wang, Yongqi, et al.
Pubblicazione: (2023)
NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation
di: Ni, Qinke, et al.
Pubblicazione: (2026)
di: Ni, Qinke, et al.
Pubblicazione: (2026)
Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding
di: Li, Bohan, et al.
Pubblicazione: (2024)
di: Li, Bohan, et al.
Pubblicazione: (2024)
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
di: Li, Haitao, et al.
Pubblicazione: (2026)
di: Li, Haitao, et al.
Pubblicazione: (2026)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
di: Han, Wooseok, et al.
Pubblicazione: (2024)
di: Han, Wooseok, et al.
Pubblicazione: (2024)
LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
di: Fang, Qingkai, et al.
Pubblicazione: (2025)
di: Fang, Qingkai, et al.
Pubblicazione: (2025)
Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
di: Liu, Zhijun, et al.
Pubblicazione: (2024)
di: Liu, Zhijun, et al.
Pubblicazione: (2024)
DSFlow: Dual Supervision and Step-Aware Architecture for One-Step Flow Matching Speech Synthesis
di: Lin, Bin, et al.
Pubblicazione: (2026)
di: Lin, Bin, et al.
Pubblicazione: (2026)
FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles
di: Zhang, Tian-Hao, et al.
Pubblicazione: (2025)
di: Zhang, Tian-Hao, et al.
Pubblicazione: (2025)
OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching
di: Huynh-Nguyen, Hieu-Nghia, et al.
Pubblicazione: (2025)
di: Huynh-Nguyen, Hieu-Nghia, et al.
Pubblicazione: (2025)
DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis
di: Li, Yingahao Aaron, et al.
Pubblicazione: (2024)
di: Li, Yingahao Aaron, et al.
Pubblicazione: (2024)
Efficient Autoregressive Audio Modeling via Next-Scale Prediction
di: Qiu, Kai, et al.
Pubblicazione: (2024)
di: Qiu, Kai, et al.
Pubblicazione: (2024)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
di: Lee, Jihwan, et al.
Pubblicazione: (2024)
di: Lee, Jihwan, et al.
Pubblicazione: (2024)
ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
di: Choi, Youngwon, et al.
Pubblicazione: (2026)
di: Choi, Youngwon, et al.
Pubblicazione: (2026)
BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching Models
di: Liang, Susan, et al.
Pubblicazione: (2025)
di: Liang, Susan, et al.
Pubblicazione: (2025)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
di: Kim, Ji-Hoon, et al.
Pubblicazione: (2024)
di: Kim, Ji-Hoon, et al.
Pubblicazione: (2024)
Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation
di: Chang, Yi, et al.
Pubblicazione: (2024)
di: Chang, Yi, et al.
Pubblicazione: (2024)
CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching
di: Zhang, Leying, et al.
Pubblicazione: (2025)
di: Zhang, Leying, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
di: Lin, Zijian, et al.
Pubblicazione: (2025) -
Unifying Speech Recognition, Synthesis and Conversion with Autoregressive Transformers
di: Cai, Runyuan, et al.
Pubblicazione: (2026) -
Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users using Intermediate ASR Features and Human Memory Models
di: Mogridge, Rhiannon, et al.
Pubblicazione: (2024) -
Beyond the Utterance: An Empirical Study of Very Long Context Speech Recognition
di: Flynn, Robert, et al.
Pubblicazione: (2026) -
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
di: Leung, Wing-Zin, et al.
Pubblicazione: (2024)