Flamed-TTS: Flow Matching Attention-Free Models for Efficient Generating and Dynamic Pacing Zero-shot Text-to-Speech
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huynh-Nguyen, Hieu-Nghia, Dang, Huynh Nguyen, Nguyen, Ngoc-Son, Nguyen, Van |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching
von: Huynh-Nguyen, Hieu-Nghia, et al.
Veröffentlicht: (2025)
von: Huynh-Nguyen, Hieu-Nghia, et al.
Veröffentlicht: (2025)
DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Factorized Discrete Flow Matching
von: Nguyen, Ngoc-Son, et al.
Veröffentlicht: (2025)
von: Nguyen, Ngoc-Son, et al.
Veröffentlicht: (2025)
DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization
von: Nguyen, Ngoc-Son, et al.
Veröffentlicht: (2026)
von: Nguyen, Ngoc-Son, et al.
Veröffentlicht: (2026)
Sing-On-Your-Beat: Simple Text-Controllable Accompaniment Generations
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2024)
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2024)
Zero-Shot Text-to-Speech for Vietnamese
von: Vu, Thi, et al.
Veröffentlicht: (2025)
von: Vu, Thi, et al.
Veröffentlicht: (2025)
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
von: Pham, The Hieu, et al.
Veröffentlicht: (2025)
von: Pham, The Hieu, et al.
Veröffentlicht: (2025)
Real-time Speech Summarization for Medical Conversations
von: Le-Duc, Khai, et al.
Veröffentlicht: (2024)
von: Le-Duc, Khai, et al.
Veröffentlicht: (2024)
Improving Pronunciation and Accent Conversion through Knowledge Distillation And Synthetic Ground-Truth from Native TTS
von: Nguyen, Tuan Nam, et al.
Veröffentlicht: (2024)
von: Nguyen, Tuan Nam, et al.
Veröffentlicht: (2024)
VietSuperSpeech: A Large-Scale Vietnamese Conversational Speech Dataset for ASR Fine-Tuning in Chatbot, Customer Support, and Call Center Applications
von: Do, Loan, et al.
Veröffentlicht: (2026)
von: Do, Loan, et al.
Veröffentlicht: (2026)
AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis
von: Luo, Dan, et al.
Veröffentlicht: (2025)
von: Luo, Dan, et al.
Veröffentlicht: (2025)
Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection
von: Truong, Duc-Tuan, et al.
Veröffentlicht: (2024)
von: Truong, Duc-Tuan, et al.
Veröffentlicht: (2024)
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
von: Deng, Wei, et al.
Veröffentlicht: (2025)
von: Deng, Wei, et al.
Veröffentlicht: (2025)
Aleatoric Uncertainty Medical Image Segmentation Estimation via Flow Matching
von: Van Nguyen, Phi, et al.
Veröffentlicht: (2025)
von: Van Nguyen, Phi, et al.
Veröffentlicht: (2025)
Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses
von: Zhao, Shengkui, et al.
Veröffentlicht: (2021)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2021)
AFSS: Artifact-Focused Self-Synthesis for Mitigating Bias in Audio Deepfake Detection
von: Nguyen-Le, Hai-Son, et al.
Veröffentlicht: (2026)
von: Nguyen-Le, Hai-Son, et al.
Veröffentlicht: (2026)
Robust TTS Training via Self-Purifying Flow Matching for the WildSpoof 2026 TTS Track
von: Yi, June Young, et al.
Veröffentlicht: (2025)
von: Yi, June Young, et al.
Veröffentlicht: (2025)
Modeling Power Systems Dynamics with Symbolic Physics-Informed Neural Networks
von: Tran, Huynh T. T., et al.
Veröffentlicht: (2023)
von: Tran, Huynh T. T., et al.
Veröffentlicht: (2023)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition
von: Anh, Tran Nguyen, et al.
Veröffentlicht: (2025)
von: Anh, Tran Nguyen, et al.
Veröffentlicht: (2025)
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2025)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2025)
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS
von: Nguyen, Tuan Nam, et al.
Veröffentlicht: (2024)
von: Nguyen, Tuan Nam, et al.
Veröffentlicht: (2024)
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
von: Wu, Zhichao, et al.
Veröffentlicht: (2025)
von: Wu, Zhichao, et al.
Veröffentlicht: (2025)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
von: Guo, Yiwei, et al.
Veröffentlicht: (2023)
von: Guo, Yiwei, et al.
Veröffentlicht: (2023)
TolerantECG: A Foundation Model for Imperfect Electrocardiogram
von: Nguyen, Huynh Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Huynh Dang, et al.
Veröffentlicht: (2025)
SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
Cocktail-Party Audio-Visual Speech Recognition
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2025)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2025)
FMSD-TTS: Few-shot Multi-Speaker Multi-Dialect Text-to-Speech Synthesis for Ü-Tsang, Amdo and Kham Speech Dataset Generation
von: Liu, Yutong, et al.
Veröffentlicht: (2025)
von: Liu, Yutong, et al.
Veröffentlicht: (2025)
Towards Lightweight and Stable Zero-shot TTS with Self-distilled Representation Disentanglement
von: Chen, Qianniu, et al.
Veröffentlicht: (2025)
von: Chen, Qianniu, et al.
Veröffentlicht: (2025)
HiFiNet: Hierarchical Fault Identification in Wireless Sensor Networks via Edge-Based Classification and Graph Aggregation
von: Nghia, Nguyen Tri, et al.
Veröffentlicht: (2025)
von: Nghia, Nguyen Tri, et al.
Veröffentlicht: (2025)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
von: Le-Duc, Khai, et al.
Veröffentlicht: (2025)
von: Le-Duc, Khai, et al.
Veröffentlicht: (2025)
Deepfake Audio Detection Using Spectrogram-based Feature and Ensemble of Deep Learning Models
von: Pham, Lam, et al.
Veröffentlicht: (2024)
von: Pham, Lam, et al.
Veröffentlicht: (2024)
LoRP-TTS: Low-Rank Personalized Text-To-Speech
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2025)
von: Bondaruk, Łukasz, et al.
Veröffentlicht: (2025)
Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models
von: Chen, Sijing, et al.
Veröffentlicht: (2024)
von: Chen, Sijing, et al.
Veröffentlicht: (2024)
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
von: Dang, Trung, et al.
Veröffentlicht: (2024)
von: Dang, Trung, et al.
Veröffentlicht: (2024)
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
von: Akti, Seymanur, et al.
Veröffentlicht: (2025)
von: Akti, Seymanur, et al.
Veröffentlicht: (2025)
A Toolchain for Comprehensive Audio/Video Analysis Using Deep Learning Based Multimodal Approach (A use case of riot or violent context detection)
von: Pham, Lam, et al.
Veröffentlicht: (2024)
von: Pham, Lam, et al.
Veröffentlicht: (2024)
EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering
von: Xie, Tianxin, et al.
Veröffentlicht: (2025)
von: Xie, Tianxin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching
von: Huynh-Nguyen, Hieu-Nghia, et al.
Veröffentlicht: (2025) -
DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Factorized Discrete Flow Matching
von: Nguyen, Ngoc-Son, et al.
Veröffentlicht: (2025) -
DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization
von: Nguyen, Ngoc-Son, et al.
Veröffentlicht: (2026) -
Sing-On-Your-Beat: Simple Text-Controllable Accompaniment Generations
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2024) -
Zero-Shot Text-to-Speech for Vietnamese
von: Vu, Thi, et al.
Veröffentlicht: (2025)