Drax: Speech Recognition with Discrete Flow Matching
Fuente:
arXiv
Salvato in:
| Autori principali: | Navon, Aviv, Shamsian, Aviv, Glazer, Neta, Segal-Feldman, Yael, Hetz, Gill, Keshet, Joseph, Fetaya, Ethan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FlowTSE: Target Speaker Extraction with Flow Matching
di: Navon, Aviv, et al.
Pubblicazione: (2025)
di: Navon, Aviv, et al.
Pubblicazione: (2025)
Keyword-Guided Adaptation of Automatic Speech Recognition
di: Shamsian, Aviv, et al.
Pubblicazione: (2024)
di: Shamsian, Aviv, et al.
Pubblicazione: (2024)
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
di: Glazer, Neta, et al.
Pubblicazione: (2025)
di: Glazer, Neta, et al.
Pubblicazione: (2025)
Beyond Transcription: Mechanistic Interpretability in ASR
di: Glazer, Neta, et al.
Pubblicazione: (2025)
di: Glazer, Neta, et al.
Pubblicazione: (2025)
Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
di: Segal-Feldman, Yael, et al.
Pubblicazione: (2024)
di: Segal-Feldman, Yael, et al.
Pubblicazione: (2024)
LipVoicer: Generating Speech from Silent Videos Guided by Lip Reading
di: Yemini, Yochai, et al.
Pubblicazione: (2023)
di: Yemini, Yochai, et al.
Pubblicazione: (2023)
Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes
di: Glazer, Neta, et al.
Pubblicazione: (2025)
di: Glazer, Neta, et al.
Pubblicazione: (2025)
Keyword Spotting with Hyper-Matched Filters for Small Footprint Devices
di: Segal-Feldman, Yael, et al.
Pubblicazione: (2025)
di: Segal-Feldman, Yael, et al.
Pubblicazione: (2025)
Multi Task Inverse Reinforcement Learning for Common Sense Reward
di: Glazer, Neta, et al.
Pubblicazione: (2024)
di: Glazer, Neta, et al.
Pubblicazione: (2024)
SSNAPS: Audio-Visual Separation of Speech and Background Noise with Diffusion Inverse Sampling
di: Yemini, Yochai, et al.
Pubblicazione: (2026)
di: Yemini, Yochai, et al.
Pubblicazione: (2026)
HebDB: a Weakly Supervised Dataset for Hebrew Speech Processing
di: Turetzky, Arnon, et al.
Pubblicazione: (2024)
di: Turetzky, Arnon, et al.
Pubblicazione: (2024)
Latent-Level Enhancement with Flow Matching for Robust Automatic Speech Recognition
di: Yang, Da-Hee, et al.
Pubblicazione: (2026)
di: Yang, Da-Hee, et al.
Pubblicazione: (2026)
RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching
di: Yang, Jinhyeok, et al.
Pubblicazione: (2026)
di: Yang, Jinhyeok, et al.
Pubblicazione: (2026)
DiffAR: Denoising Diffusion Autoregressive Model for Raw Speech Waveform Generation
di: Benita, Roi, et al.
Pubblicazione: (2023)
di: Benita, Roi, et al.
Pubblicazione: (2023)
WhisperNER: Unified Open Named Entity and Speech Recognition
di: Ayache, Gil, et al.
Pubblicazione: (2024)
di: Ayache, Gil, et al.
Pubblicazione: (2024)
Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
di: Wang, Huimeng, et al.
Pubblicazione: (2025)
di: Wang, Huimeng, et al.
Pubblicazione: (2025)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
di: Chen, Peikun, et al.
Pubblicazione: (2024)
di: Chen, Peikun, et al.
Pubblicazione: (2024)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
di: Jung, Chaeyoung, et al.
Pubblicazione: (2024)
di: Jung, Chaeyoung, et al.
Pubblicazione: (2024)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
di: Wang, Shih-heng, et al.
Pubblicazione: (2024)
di: Wang, Shih-heng, et al.
Pubblicazione: (2024)
FlowW2N: Whispered-to-Normal Speech Conversion via Flow-Matching
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2026)
di: Ritter-Gutierrez, Fabian, et al.
Pubblicazione: (2026)
On-device Streaming Discrete Speech Units
di: Choi, Kwanghee, et al.
Pubblicazione: (2025)
di: Choi, Kwanghee, et al.
Pubblicazione: (2025)
Recovering Performance in Speech Emotion Recognition from Discrete Tokens via Multi-Layer Fusion and Paralinguistic Feature Integration
di: Sun, Esther, et al.
Pubblicazione: (2026)
di: Sun, Esther, et al.
Pubblicazione: (2026)
Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model
di: Zuo, Jialong, et al.
Pubblicazione: (2025)
di: Zuo, Jialong, et al.
Pubblicazione: (2025)
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
di: Xie, Hanke, et al.
Pubblicazione: (2025)
di: Xie, Hanke, et al.
Pubblicazione: (2025)
VoiceRestore: Flow-Matching Transformers for Speech Recording Quality Restoration
di: Kirdey, Stanislav
Pubblicazione: (2025)
di: Kirdey, Stanislav
Pubblicazione: (2025)
WhisperRT -- Turning Whisper into a Causal Streaming Model
di: Krichli, Tomer, et al.
Pubblicazione: (2025)
di: Krichli, Tomer, et al.
Pubblicazione: (2025)
StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
di: Guo, Dake, et al.
Pubblicazione: (2025)
di: Guo, Dake, et al.
Pubblicazione: (2025)
PatchDSU: Uncertainty Modeling for Out of Distribution Generalization in Keyword Spotting
di: Chernyak, Bronya Roni, et al.
Pubblicazione: (2025)
di: Chernyak, Bronya Roni, et al.
Pubblicazione: (2025)
Multi-blank Transducers for Speech Recognition
di: Xu, Hainan, et al.
Pubblicazione: (2022)
di: Xu, Hainan, et al.
Pubblicazione: (2022)
Generative Pre-training for Speech with Flow Matching
di: Liu, Alexander H., et al.
Pubblicazione: (2023)
di: Liu, Alexander H., et al.
Pubblicazione: (2023)
Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
di: Zheng, Qixi, et al.
Pubblicazione: (2025)
di: Zheng, Qixi, et al.
Pubblicazione: (2025)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
di: Zhu, Han, et al.
Pubblicazione: (2025)
di: Zhu, Han, et al.
Pubblicazione: (2025)
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
di: Chen, Yushen, et al.
Pubblicazione: (2024)
di: Chen, Yushen, et al.
Pubblicazione: (2024)
High-Fidelity Speech Enhancement via Discrete Audio Tokens
di: Lanzendörfer, Luca A., et al.
Pubblicazione: (2025)
di: Lanzendörfer, Luca A., et al.
Pubblicazione: (2025)
Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis
di: Jiang, Xilin, et al.
Pubblicazione: (2024)
di: Jiang, Xilin, et al.
Pubblicazione: (2024)
Test-Time Adaptation for Speech Emotion Recognition
di: Dong, Jiaheng, et al.
Pubblicazione: (2026)
di: Dong, Jiaheng, et al.
Pubblicazione: (2026)
Adapting WavLM for Speech Emotion Recognition
di: Diatlova, Daria, et al.
Pubblicazione: (2024)
di: Diatlova, Daria, et al.
Pubblicazione: (2024)
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech
di: Zhang, Xu, et al.
Pubblicazione: (2026)
di: Zhang, Xu, et al.
Pubblicazione: (2026)
In-Materia Speech Recognition
di: Zolfagharinejad, Mohamadreza, et al.
Pubblicazione: (2024)
di: Zolfagharinejad, Mohamadreza, et al.
Pubblicazione: (2024)
Absorbing Discrete Diffusion for Speech Enhancement
di: Gonzalez, Philippe
Pubblicazione: (2026)
di: Gonzalez, Philippe
Pubblicazione: (2026)
Documenti analoghi
-
FlowTSE: Target Speaker Extraction with Flow Matching
di: Navon, Aviv, et al.
Pubblicazione: (2025) -
Keyword-Guided Adaptation of Automatic Speech Recognition
di: Shamsian, Aviv, et al.
Pubblicazione: (2024) -
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
di: Glazer, Neta, et al.
Pubblicazione: (2025) -
Beyond Transcription: Mechanistic Interpretability in ASR
di: Glazer, Neta, et al.
Pubblicazione: (2025) -
Whisper in Medusa's Ear: Multi-head Efficient Decoding for Transformer-based ASR
di: Segal-Feldman, Yael, et al.
Pubblicazione: (2024)