FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Hui, Liu, Shujie, Meng, Lingwei, Li, Jinyu, Yang, Yifan, Zhao, Shiwan, Sun, Haiyang, Liu, Yanqing, Sun, Haoqin, Zhou, Jiaming, Lu, Yan, Qin, Yong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
di: Wang, Hui, et al.
Pubblicazione: (2025)
di: Wang, Hui, et al.
Pubblicazione: (2025)
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
di: Wang, Hui, et al.
Pubblicazione: (2025)
di: Wang, Hui, et al.
Pubblicazione: (2025)
Autoregressive Speech Synthesis without Vector Quantization
di: Meng, Lingwei, et al.
Pubblicazione: (2024)
di: Meng, Lingwei, et al.
Pubblicazione: (2024)
Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
di: Sun, Haoqin, et al.
Pubblicazione: (2024)
di: Sun, Haoqin, et al.
Pubblicazione: (2024)
WildElder: A Chinese Elderly Speech Dataset from the Wild with Fine-Grained Manual Annotations
di: Wang, Hui, et al.
Pubblicazione: (2025)
di: Wang, Hui, et al.
Pubblicazione: (2025)
Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
di: Yang, Dong, et al.
Pubblicazione: (2025)
di: Yang, Dong, et al.
Pubblicazione: (2025)
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
di: Yuan, Ze, et al.
Pubblicazione: (2024)
di: Yuan, Ze, et al.
Pubblicazione: (2024)
A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
di: Wang, Shiyao, et al.
Pubblicazione: (2025)
di: Wang, Shiyao, et al.
Pubblicazione: (2025)
VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
di: Han, Bing, et al.
Pubblicazione: (2024)
di: Han, Bing, et al.
Pubblicazione: (2024)
Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
di: Wang, Xuechen, et al.
Pubblicazione: (2024)
di: Wang, Xuechen, et al.
Pubblicazione: (2024)
Next Tokens Denoising for Speech Synthesis
di: Liu, Yanqing, et al.
Pubblicazione: (2025)
di: Liu, Yanqing, et al.
Pubblicazione: (2025)
MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation
di: Liu, Cheng, et al.
Pubblicazione: (2025)
di: Liu, Cheng, et al.
Pubblicazione: (2025)
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
di: Sun, Haoqin, et al.
Pubblicazione: (2025)
di: Sun, Haoqin, et al.
Pubblicazione: (2025)
TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models
di: Wang, Hui, et al.
Pubblicazione: (2025)
di: Wang, Hui, et al.
Pubblicazione: (2025)
ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5
di: Zhou, Jiaming, et al.
Pubblicazione: (2024)
di: Zhou, Jiaming, et al.
Pubblicazione: (2024)
M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
di: Zhou, Jiaming, et al.
Pubblicazione: (2024)
di: Zhou, Jiaming, et al.
Pubblicazione: (2024)
Uncertainty-Aware Mean Opinion Score Prediction
di: Wang, Hui, et al.
Pubblicazione: (2024)
di: Wang, Hui, et al.
Pubblicazione: (2024)
Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
di: Zhou, Jiaming, et al.
Pubblicazione: (2024)
di: Zhou, Jiaming, et al.
Pubblicazione: (2024)
Parallel Synthesis for Autoregressive Speech Generation
di: Hsu, Po-chun, et al.
Pubblicazione: (2022)
di: Hsu, Po-chun, et al.
Pubblicazione: (2022)
Boosting Large Language Model for Speech Synthesis: An Empirical Study
di: Hao, Hongkun, et al.
Pubblicazione: (2023)
di: Hao, Hongkun, et al.
Pubblicazione: (2023)
PB-LRDWWS System for the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting Challenge
di: Wang, Shiyao, et al.
Pubblicazione: (2024)
di: Wang, Shiyao, et al.
Pubblicazione: (2024)
kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
di: Zhou, Jiaming, et al.
Pubblicazione: (2023)
di: Zhou, Jiaming, et al.
Pubblicazione: (2023)
Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation
di: Wang, Shiyao, et al.
Pubblicazione: (2024)
di: Wang, Shiyao, et al.
Pubblicazione: (2024)
DIFFA: Large Language Diffusion Models Can Listen and Understand
di: Zhou, Jiaming, et al.
Pubblicazione: (2025)
di: Zhou, Jiaming, et al.
Pubblicazione: (2025)
CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English Code-Switching Dialogues for Speech Recognition
di: Zhou, Jiaming, et al.
Pubblicazione: (2025)
di: Zhou, Jiaming, et al.
Pubblicazione: (2025)
AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
di: Wang, Hui, et al.
Pubblicazione: (2025)
di: Wang, Hui, et al.
Pubblicazione: (2025)
Advanced Long-Content Speech Recognition With Factorized Neural Transducer
di: Gong, Xun, et al.
Pubblicazione: (2024)
di: Gong, Xun, et al.
Pubblicazione: (2024)
A Unified Neural Codec Language Model for Selective Editable Text to Speech Generation
di: Pei, Hanchen, et al.
Pubblicazione: (2026)
di: Pei, Hanchen, et al.
Pubblicazione: (2026)
Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
di: Yang, Yifan, et al.
Pubblicazione: (2025)
di: Yang, Yifan, et al.
Pubblicazione: (2025)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
di: Li, Jiaqi, et al.
Pubblicazione: (2024)
di: Li, Jiaqi, et al.
Pubblicazione: (2024)
StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
di: Guo, Dake, et al.
Pubblicazione: (2025)
di: Guo, Dake, et al.
Pubblicazione: (2025)
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
di: Wang, Huimeng, et al.
Pubblicazione: (2025)
di: Wang, Huimeng, et al.
Pubblicazione: (2025)
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
di: Chen, Sanyuan, et al.
Pubblicazione: (2024)
di: Chen, Sanyuan, et al.
Pubblicazione: (2024)
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
di: Choi, Jeongsoo, et al.
Pubblicazione: (2024)
di: Choi, Jeongsoo, et al.
Pubblicazione: (2024)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
di: Pham, The Hieu, et al.
Pubblicazione: (2025)
di: Pham, The Hieu, et al.
Pubblicazione: (2025)
WavLLM: Towards Robust and Adaptive Speech Large Language Model
di: Hu, Shujie, et al.
Pubblicazione: (2024)
di: Hu, Shujie, et al.
Pubblicazione: (2024)
Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
di: Wang, Huimeng, et al.
Pubblicazione: (2024)
di: Wang, Huimeng, et al.
Pubblicazione: (2024)
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
di: Xie, Hanke, et al.
Pubblicazione: (2025)
di: Xie, Hanke, et al.
Pubblicazione: (2025)
Autoregressive Speech Enhancement via Acoustic Tokens
di: Della Libera, Luca, et al.
Pubblicazione: (2025)
di: Della Libera, Luca, et al.
Pubblicazione: (2025)
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
di: Lin, Zijian, et al.
Pubblicazione: (2025)
di: Lin, Zijian, et al.
Pubblicazione: (2025)
Documenti analoghi
-
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
di: Wang, Hui, et al.
Pubblicazione: (2025) -
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
di: Wang, Hui, et al.
Pubblicazione: (2025) -
Autoregressive Speech Synthesis without Vector Quantization
di: Meng, Lingwei, et al.
Pubblicazione: (2024) -
Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
di: Sun, Haoqin, et al.
Pubblicazione: (2024) -
WildElder: A Chinese Elderly Speech Dataset from the Wild with Fine-Grained Manual Annotations
di: Wang, Hui, et al.
Pubblicazione: (2025)