FlowW2N: Whispered-to-Normal Speech Conversion via Flow-Matching
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ritter-Gutierrez, Fabian, Jalal, Md Asif, Parada, Pablo Peso, Saravanan, Karthikeyan, Shul, Yusun, Kim, Minseung, Lee, Gun-Woo, Moon, Han-Gil |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning
von: Zhang, Jisi, et al.
Veröffentlicht: (2025)
von: Zhang, Jisi, et al.
Veröffentlicht: (2025)
CST-former: Multidimensional Attention-based Transformer for Sound Event Localization and Detection in Real Scenes
von: Shul, Yusun, et al.
Veröffentlicht: (2025)
von: Shul, Yusun, et al.
Veröffentlicht: (2025)
persoDA: Personalized Data Augmentation for Personalized ASR
von: Parada, Pablo Peso, et al.
Veröffentlicht: (2025)
von: Parada, Pablo Peso, et al.
Veröffentlicht: (2025)
Consistency Based Unsupervised Self-training For ASR Personalisation
von: Zhang, Jisi, et al.
Veröffentlicht: (2024)
von: Zhang, Jisi, et al.
Veröffentlicht: (2024)
Exploring compressibility of transformer based text-to-music (TTM) models
von: Moschopoulos, Vasileios, et al.
Veröffentlicht: (2024)
von: Moschopoulos, Vasileios, et al.
Veröffentlicht: (2024)
Locality enhanced dynamic biasing and sampling strategies for contextual ASR
von: Jalal, Md Asif, et al.
Veröffentlicht: (2024)
von: Jalal, Md Asif, et al.
Veröffentlicht: (2024)
Retrieval Augmented Generation based context discovery for ASR
von: Siskos, Dimitrios, et al.
Veröffentlicht: (2025)
von: Siskos, Dimitrios, et al.
Veröffentlicht: (2025)
ValSub: Subsampling Validation Data to Mitigate Forgetting during ASR Personalization
von: Mehmood, Haaris, et al.
Veröffentlicht: (2025)
von: Mehmood, Haaris, et al.
Veröffentlicht: (2025)
WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
MaskCycleGAN-based Whisper to Normal Speech Conversion
von: Gupta, K. Rohith, et al.
Veröffentlicht: (2024)
von: Gupta, K. Rohith, et al.
Veröffentlicht: (2024)
FlowSE: Flow Matching-based Speech Enhancement
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
von: Jung, Chaeyoung, et al.
Veröffentlicht: (2024)
FlowSE-GRPO: Training Flow Matching Speech Enhancement via Online Reinforcement Learning
von: Wang, Haoxu, et al.
Veröffentlicht: (2026)
von: Wang, Haoxu, et al.
Veröffentlicht: (2026)
Towards Real-Time Generative Speech Restoration with Flow-Matching
von: Hsieh, Tsun-An, et al.
Veröffentlicht: (2025)
von: Hsieh, Tsun-An, et al.
Veröffentlicht: (2025)
WhispEar: A Bi-directional Framework for Scaling Whispered Speech Conversion via Pseudo-Parallel Whisper Generation
von: Fang, Zihao, et al.
Veröffentlicht: (2026)
von: Fang, Zihao, et al.
Veröffentlicht: (2026)
WhisperFlow: speech foundation models in real time
von: Wang, Rongxiang, et al.
Veröffentlicht: (2024)
von: Wang, Rongxiang, et al.
Veröffentlicht: (2024)
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
Drax: Speech Recognition with Discrete Flow Matching
von: Navon, Aviv, et al.
Veröffentlicht: (2025)
von: Navon, Aviv, et al.
Veröffentlicht: (2025)
FlowTSE: Target Speaker Extraction with Flow Matching
von: Navon, Aviv, et al.
Veröffentlicht: (2025)
von: Navon, Aviv, et al.
Veröffentlicht: (2025)
Quartered Chirp Spectral Envelope for Whispered vs Normal Speech Classification
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
von: Guo, Dake, et al.
Veröffentlicht: (2025)
von: Guo, Dake, et al.
Veröffentlicht: (2025)
LP-CFM: Perceptual Invariance-Aware Conditional Flow Matching for Speech Modeling
von: Kwak, Doyeop, et al.
Veröffentlicht: (2025)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2025)
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
RapFlow-TTS: Rapid and High-Fidelity Text-to-Speech with Improved Consistency Flow Matching
von: Park, Hyun Joon, et al.
Veröffentlicht: (2025)
von: Park, Hyun Joon, et al.
Veröffentlicht: (2025)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
von: Guo, Yiwei, et al.
Veröffentlicht: (2023)
von: Guo, Yiwei, et al.
Veröffentlicht: (2023)
RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2026)
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2026)
Latent-Level Enhancement with Flow Matching for Robust Automatic Speech Recognition
von: Yang, Da-Hee, et al.
Veröffentlicht: (2026)
von: Yang, Da-Hee, et al.
Veröffentlicht: (2026)
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
VoiceRestore: Flow-Matching Transformers for Speech Recording Quality Restoration
von: Kirdey, Stanislav
Veröffentlicht: (2025)
von: Kirdey, Stanislav
Veröffentlicht: (2025)
Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech
von: Avdeeva, Anastasia, et al.
Veröffentlicht: (2024)
von: Avdeeva, Anastasia, et al.
Veröffentlicht: (2024)
Generative Pre-training for Speech with Flow Matching
von: Liu, Alexander H., et al.
Veröffentlicht: (2023)
von: Liu, Alexander H., et al.
Veröffentlicht: (2023)
Quartered Spectral Envelope and 1D-CNN-based Classification of Normally Phonated and Whispered Speech
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
von: Joysingh, S. Johanan, et al.
Veröffentlicht: (2024)
Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
von: Yang, Dong, et al.
Veröffentlicht: (2025)
von: Yang, Dong, et al.
Veröffentlicht: (2025)
Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching
von: Das, Shoutrik, et al.
Veröffentlicht: (2025)
von: Das, Shoutrik, et al.
Veröffentlicht: (2025)
Vocoder-Free Non-Parallel Conversion of Whispered Speech With Masked Cycle-Consistent Generative Adversarial Networks
von: Wagner, Dominik, et al.
Veröffentlicht: (2023)
von: Wagner, Dominik, et al.
Veröffentlicht: (2023)
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
von: Zuo, Jialong, et al.
Veröffentlicht: (2025)
LatentVoiceGrad: Nonparallel Voice Conversion with Latent Diffusion/Flow-Matching Models
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2025)
von: Kameoka, Hirokazu, et al.
Veröffentlicht: (2025)
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning
von: Zhang, Jisi, et al.
Veröffentlicht: (2025) -
CST-former: Multidimensional Attention-based Transformer for Sound Event Localization and Detection in Real Scenes
von: Shul, Yusun, et al.
Veröffentlicht: (2025) -
persoDA: Personalized Data Augmentation for Personalized ASR
von: Parada, Pablo Peso, et al.
Veröffentlicht: (2025) -
Consistency Based Unsupervised Self-training For ASR Personalisation
von: Zhang, Jisi, et al.
Veröffentlicht: (2024) -
Exploring compressibility of transformer based text-to-music (TTM) models
von: Moschopoulos, Vasileios, et al.
Veröffentlicht: (2024)