SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Kaidi, Guan, Wenhao, Lu, Shenghui, Yao, Jianglong, Li, Lin, Hong, Qingyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement
von: Lu, Shenghui, et al.
Veröffentlicht: (2025)
von: Lu, Shenghui, et al.
Veröffentlicht: (2025)
ReFlow-TTS: A Rectified Flow Model for High-fidelity Text-to-Speech
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
von: Guo, Yiwei, et al.
Veröffentlicht: (2023)
von: Guo, Yiwei, et al.
Veröffentlicht: (2023)
MeanFlowSE: one-step generative speech enhancement via conditional mean flow
von: Li, Duojia, et al.
Veröffentlicht: (2025)
von: Li, Duojia, et al.
Veröffentlicht: (2025)
ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
Adaptive Slimming for Scalable and Efficient Speech Enhancement
von: Miccini, Riccardo, et al.
Veröffentlicht: (2025)
von: Miccini, Riccardo, et al.
Veröffentlicht: (2025)
Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
RFM-Editing: Rectified Flow Matching for Text-guided Audio Editing
von: Gao, Liting, et al.
Veröffentlicht: (2025)
von: Gao, Liting, et al.
Veröffentlicht: (2025)
Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model
von: Huang, Hukai, et al.
Veröffentlicht: (2024)
von: Huang, Hukai, et al.
Veröffentlicht: (2024)
Pseudo Labels-based Neural Speech Enhancement for the AVSR Task in the MISP-Meeting Challenge
von: Luo, Longjie, et al.
Veröffentlicht: (2025)
von: Luo, Longjie, et al.
Veröffentlicht: (2025)
DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec
von: Chen, Peijie, et al.
Veröffentlicht: (2025)
von: Chen, Peijie, et al.
Veröffentlicht: (2025)
Flamed-TTS: Flow Matching Attention-Free Models for Efficient Generating and Dynamic Pacing Zero-shot Text-to-Speech
von: Huynh-Nguyen, Hieu-Nghia, et al.
Veröffentlicht: (2025)
von: Huynh-Nguyen, Hieu-Nghia, et al.
Veröffentlicht: (2025)
Sample-Efficient Diffusion for Text-To-Speech Synthesis
von: Lovelace, Justin, et al.
Veröffentlicht: (2024)
von: Lovelace, Justin, et al.
Veröffentlicht: (2024)
AlphaFlowTSE: One-Step Generative Target Speaker Extraction via Conditional AlphaFlow
von: Li, Duojia, et al.
Veröffentlicht: (2026)
von: Li, Duojia, et al.
Veröffentlicht: (2026)
Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
von: Yang, Dong, et al.
Veröffentlicht: (2025)
von: Yang, Dong, et al.
Veröffentlicht: (2025)
LSZone: A Lightweight Spatial Information Modeling Architecture for Real-time In-car Multi-zone Speech Separation
von: Chen, Jun, et al.
Veröffentlicht: (2025)
von: Chen, Jun, et al.
Veröffentlicht: (2025)
SLM-SS: Speech Language Model for Generative Speech Separation
von: Li, Tianhua, et al.
Veröffentlicht: (2026)
von: Li, Tianhua, et al.
Veröffentlicht: (2026)
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
von: Deng, Wei, et al.
Veröffentlicht: (2025)
von: Deng, Wei, et al.
Veröffentlicht: (2025)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
Fake Speech Wild: Detecting Deepfake Speech on Social Media Platform
von: Xie, Yuankun, et al.
Veröffentlicht: (2025)
von: Xie, Yuankun, et al.
Veröffentlicht: (2025)
Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction
von: Wu, Weijie, et al.
Veröffentlicht: (2025)
von: Wu, Weijie, et al.
Veröffentlicht: (2025)
A Lightweight Pipeline for Noisy Speech Voice Cloning and Accurate Lip Sync Synthesis
von: Amir, Javeria, et al.
Veröffentlicht: (2025)
von: Amir, Javeria, et al.
Veröffentlicht: (2025)
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
von: Li, Haitao, et al.
Veröffentlicht: (2026)
von: Li, Haitao, et al.
Veröffentlicht: (2026)
DDSP-QbE++: Improving Speech Quality for Speech Anonymisation for Atypical Speech
von: Ghosh, Suhita, et al.
Veröffentlicht: (2026)
von: Ghosh, Suhita, et al.
Veröffentlicht: (2026)
Soundwave: Less is More for Speech-Text Alignment in LLMs
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
GSRM: Generative Speech Reward Model for Speech RLHF
von: Shen, Maohao, et al.
Veröffentlicht: (2026)
von: Shen, Maohao, et al.
Veröffentlicht: (2026)
Emotion Detection in Speech Using Lightweight and Transformer-Based Models: A Comparative and Ablation Study
von: Onyekwelu-Udoka, Lucky, et al.
Veröffentlicht: (2025)
von: Onyekwelu-Udoka, Lucky, et al.
Veröffentlicht: (2025)
StyleSpeech: Parameter-efficient Fine Tuning for Pre-trained Controllable Text-to-Speech
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
SepPrune: Structured Pruning for Efficient Deep Speech Separation
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
Efficient Training for Cross-lingual Speech Language Models
von: Zhou, Yan, et al.
Veröffentlicht: (2026)
von: Zhou, Yan, et al.
Veröffentlicht: (2026)
AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis
von: Luo, Dan, et al.
Veröffentlicht: (2025)
von: Luo, Dan, et al.
Veröffentlicht: (2025)
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
von: Wang, Helin, et al.
Veröffentlicht: (2025)
von: Wang, Helin, et al.
Veröffentlicht: (2025)
LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
von: Guan, Wenhao, et al.
Veröffentlicht: (2024)
von: Guan, Wenhao, et al.
Veröffentlicht: (2024)
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer
von: Wang, Yongqi, et al.
Veröffentlicht: (2023)
von: Wang, Yongqi, et al.
Veröffentlicht: (2023)
Generalizable Speech Deepfake Detection via Information Bottleneck Enhanced Adversarial Alignment
von: Huang, Pu, et al.
Veröffentlicht: (2025)
von: Huang, Pu, et al.
Veröffentlicht: (2025)
SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
Understanding Frechet Speech Distance for Synthetic Speech Quality Evaluation
von: Kim, June-Woo, et al.
Veröffentlicht: (2026)
von: Kim, June-Woo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement
von: Lu, Shenghui, et al.
Veröffentlicht: (2025) -
ReFlow-TTS: A Rectified Flow Model for High-fidelity Text-to-Speech
von: Guan, Wenhao, et al.
Veröffentlicht: (2023) -
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
von: Guo, Yiwei, et al.
Veröffentlicht: (2023) -
MeanFlowSE: one-step generative speech enhancement via conditional mean flow
von: Li, Duojia, et al.
Veröffentlicht: (2025) -
ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)