VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Yiwei, Du, Chenpeng, Ma, Ziyang, Chen, Xie, Yu, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
by: Du, Chenpeng, et al.
Published: (2022)
by: Du, Chenpeng, et al.
Published: (2022)
VoiceX: A Text-To-Speech Framework for Custom Voices
by: Mertes, Silvan, et al.
Published: (2024)
by: Mertes, Silvan, et al.
Published: (2024)
Language Model Can Listen While Speaking
by: Ma, Ziyang, et al.
Published: (2024)
by: Ma, Ziyang, et al.
Published: (2024)
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
by: Wang, Hankun, et al.
Published: (2024)
by: Wang, Hankun, et al.
Published: (2024)
Acoustic BPE for Speech Generation with Discrete Tokens
by: Shen, Feiyu, et al.
Published: (2023)
by: Shen, Feiyu, et al.
Published: (2023)
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
by: Guo, Yiwei, et al.
Published: (2024)
by: Guo, Yiwei, et al.
Published: (2024)
Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
by: Zheng, Qixi, et al.
Published: (2025)
by: Zheng, Qixi, et al.
Published: (2025)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
by: Wang, Hongbin, et al.
Published: (2025)
by: Wang, Hongbin, et al.
Published: (2025)
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
by: Guo, Yiwei, et al.
Published: (2024)
by: Guo, Yiwei, et al.
Published: (2024)
Advancing User-Voice Interaction: Exploring Emotion-Aware Voice Assistants Through a Role-Swapping Approach
by: Ma, Yong, et al.
Published: (2025)
by: Ma, Yong, et al.
Published: (2025)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
by: Du, Chenpeng, et al.
Published: (2024)
by: Du, Chenpeng, et al.
Published: (2024)
Towards Reliable Large Audio Language Model
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
by: Du, Chenpeng, et al.
Published: (2021)
by: Du, Chenpeng, et al.
Published: (2021)
CabinSep: IR-Augmented Mask-Based MVDR for Real-Time In-Car Speech Separation with Distributed Heterogeneous Arrays
by: Han, Runduo, et al.
Published: (2025)
by: Han, Runduo, et al.
Published: (2025)
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
by: Chen, Yushen, et al.
Published: (2024)
by: Chen, Yushen, et al.
Published: (2024)
Transhuman Ansambl - Voice Beyond Language
by: Ivsic, Lucija, et al.
Published: (2024)
by: Ivsic, Lucija, et al.
Published: (2024)
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
by: Du, Chenpeng, et al.
Published: (2023)
by: Du, Chenpeng, et al.
Published: (2023)
ReFlow-TTS: A Rectified Flow Model for High-fidelity Text-to-Speech
by: Guan, Wenhao, et al.
Published: (2023)
by: Guan, Wenhao, et al.
Published: (2023)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
by: Chen, Youjun, et al.
Published: (2025)
by: Chen, Youjun, et al.
Published: (2025)
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
by: Dutta, Satwik, et al.
Published: (2025)
by: Dutta, Satwik, et al.
Published: (2025)
Multi-Speaker Multi-Lingual VQTTS System for LIMMITS 2023 Challenge
by: Du, Chenpeng, et al.
Published: (2023)
by: Du, Chenpeng, et al.
Published: (2023)
SingVisio: Visual Analytics of Diffusion Model for Singing Voice Conversion
by: Xue, Liumeng, et al.
Published: (2024)
by: Xue, Liumeng, et al.
Published: (2024)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
by: Zhu, Han, et al.
Published: (2025)
by: Zhu, Han, et al.
Published: (2025)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
by: Khanday, Owais Mujtaba, et al.
Published: (2025)
by: Khanday, Owais Mujtaba, et al.
Published: (2025)
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
by: Yu, Luca Jiang-Tao, et al.
Published: (2024)
by: Yu, Luca Jiang-Tao, et al.
Published: (2024)
Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness
by: Yang, Sicheng, et al.
Published: (2024)
by: Yang, Sicheng, et al.
Published: (2024)
A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
by: Benster, Tyler, et al.
Published: (2024)
by: Benster, Tyler, et al.
Published: (2024)
Beyond-Voice: Towards Continuous 3D Hand Pose Tracking on Commercial Home Assistant Devices
by: Li, Yin, et al.
Published: (2023)
by: Li, Yin, et al.
Published: (2023)
STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition
by: Chang, Yi, et al.
Published: (2024)
by: Chang, Yi, et al.
Published: (2024)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
by: Li, Yue, et al.
Published: (2024)
by: Li, Yue, et al.
Published: (2024)
Springboard, Roadblock or "Crutch"?: How Transgender Users Leverage Voice Changers for Gender Presentation in Social Virtual Reality
by: Povinelli, Kassie, et al.
Published: (2024)
by: Povinelli, Kassie, et al.
Published: (2024)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
by: Park, Seohyun, et al.
Published: (2025)
by: Park, Seohyun, et al.
Published: (2025)
G-STAR: End-to-End Global Speaker-Tracking Attributed Recognition
by: Peng, Jing, et al.
Published: (2026)
by: Peng, Jing, et al.
Published: (2026)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
by: Zhou, Dongliang, et al.
Published: (2025)
by: Zhou, Dongliang, et al.
Published: (2025)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
by: Feng, Tiantian, et al.
Published: (2023)
by: Feng, Tiantian, et al.
Published: (2023)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
by: Xiao, Yi, et al.
Published: (2022)
by: Xiao, Yi, et al.
Published: (2022)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
by: Cai, Zhuojiang, et al.
Published: (2024)
by: Cai, Zhuojiang, et al.
Published: (2024)
NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural Activity
by: Khanday, Owais Mujtaba, et al.
Published: (2025)
by: Khanday, Owais Mujtaba, et al.
Published: (2025)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
by: Nowrin, Sadia, et al.
Published: (2024)
by: Nowrin, Sadia, et al.
Published: (2024)
SoundShift: Exploring Sound Manipulations for Accessible Mixed-Reality Awareness
by: Chang, Ruei-Che, et al.
Published: (2024)
by: Chang, Ruei-Che, et al.
Published: (2024)
Similar Items
-
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
by: Du, Chenpeng, et al.
Published: (2022) -
VoiceX: A Text-To-Speech Framework for Custom Voices
by: Mertes, Silvan, et al.
Published: (2024) -
Language Model Can Listen While Speaking
by: Ma, Ziyang, et al.
Published: (2024) -
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
by: Wang, Hankun, et al.
Published: (2024) -
Acoustic BPE for Speech Generation with Discrete Tokens
by: Shen, Feiyu, et al.
Published: (2023)