RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Jinhyeok, Kim, Hyeongju, Yu, Yechan, Byun, Joon, Bous, Frederik, Lee, Juheon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
Training Flow Matching Models with Reliable Labels via Self-Purification
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
Length-Aware Rotary Position Embedding for Text-Speech Alignment
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2024)
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2024)
Improving Test-Time Performance of RVQ-based Neural Codecs
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)
Latent-Level Enhancement with Flow Matching for Robust Automatic Speech Recognition
von: Yang, Da-Hee, et al.
Veröffentlicht: (2026)
von: Yang, Da-Hee, et al.
Veröffentlicht: (2026)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
FlowW2N: Whispered-to-Normal Speech Conversion via Flow-Matching
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2026)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2026)
Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
von: Zheng, Qixi, et al.
Veröffentlicht: (2025)
von: Zheng, Qixi, et al.
Veröffentlicht: (2025)
Few-step Adversarial Schrödinger Bridge for Generative Speech Enhancement
von: Han, Seungu, et al.
Veröffentlicht: (2025)
von: Han, Seungu, et al.
Veröffentlicht: (2025)
FLOWER: Flow-Based Estimated Gaussian Guidance for General Speech Restoration
von: Yang, Da-Hee, et al.
Veröffentlicht: (2025)
von: Yang, Da-Hee, et al.
Veröffentlicht: (2025)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
von: Zhu, Han, et al.
Veröffentlicht: (2025)
von: Zhu, Han, et al.
Veröffentlicht: (2025)
DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers
von: Cao, Tianyu, et al.
Veröffentlicht: (2026)
von: Cao, Tianyu, et al.
Veröffentlicht: (2026)
F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization
von: Sun, Xiaohui, et al.
Veröffentlicht: (2025)
von: Sun, Xiaohui, et al.
Veröffentlicht: (2025)
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
von: Yang, Dong, et al.
Veröffentlicht: (2025)
von: Yang, Dong, et al.
Veröffentlicht: (2025)
ReFlow-TTS: A Rectified Flow Model for High-fidelity Text-to-Speech
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
VoiceRestore: Flow-Matching Transformers for Speech Recording Quality Restoration
von: Kirdey, Stanislav
Veröffentlicht: (2025)
von: Kirdey, Stanislav
Veröffentlicht: (2025)
StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
von: Guo, Dake, et al.
Veröffentlicht: (2025)
von: Guo, Dake, et al.
Veröffentlicht: (2025)
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
von: Chen, Yushen, et al.
Veröffentlicht: (2024)
von: Chen, Yushen, et al.
Veröffentlicht: (2024)
MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement
von: Nasretdinov, Rauf, et al.
Veröffentlicht: (2025)
von: Nasretdinov, Rauf, et al.
Veröffentlicht: (2025)
A Neural Speech Codec for Noise Robust Speech Coding
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
Drax: Speech Recognition with Discrete Flow Matching
von: Navon, Aviv, et al.
Veröffentlicht: (2025)
von: Navon, Aviv, et al.
Veröffentlicht: (2025)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
von: Leung, Wing-Zin, et al.
Veröffentlicht: (2024)
von: Leung, Wing-Zin, et al.
Veröffentlicht: (2024)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
Robust One-step Speech Enhancement via Consistency Distillation
von: Xu, Liang, et al.
Veröffentlicht: (2025)
von: Xu, Liang, et al.
Veröffentlicht: (2025)
Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation
von: Han, Changjin, et al.
Veröffentlicht: (2024)
von: Han, Changjin, et al.
Veröffentlicht: (2024)
Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
von: Chae, Yunkee, et al.
Veröffentlicht: (2025)
von: Chae, Yunkee, et al.
Veröffentlicht: (2025)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
von: Guo, Yiwei, et al.
Veröffentlicht: (2023)
von: Guo, Yiwei, et al.
Veröffentlicht: (2023)
Probing the Robustness Properties of Neural Speech Codecs
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
von: Le, Khanh, et al.
Veröffentlicht: (2025)
von: Le, Khanh, et al.
Veröffentlicht: (2025)
Emotion-Coherent Speech Data Augmentation and Self-Supervised Contrastive Style Training for Enhancing Kids's Story Speech Synthesis
von: Chung, Raymond
Veröffentlicht: (2026)
von: Chung, Raymond
Veröffentlicht: (2026)
Ähnliche Einträge
-
SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025) -
Training Flow Matching Models with Reliable Labels via Self-Purification
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025) -
Length-Aware Rotary Position Embedding for Text-Speech Alignment
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025) -
DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2024) -
Improving Test-Time Performance of RVQ-based Neural Codecs
von: Kim, Hyeongju, et al.
Veröffentlicht: (2025)