Real-time Speech Restoration using Data Prediction Mean Flows
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Braun, Sebastian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Real-Time Generative Speech Restoration with Flow-Matching
von: Hsieh, Tsun-An, et al.
Veröffentlicht: (2025)
von: Hsieh, Tsun-An, et al.
Veröffentlicht: (2025)
MeanSE: Efficient Generative Speech Enhancement with Mean Flows
von: Wang, Jiahe, et al.
Veröffentlicht: (2025)
von: Wang, Jiahe, et al.
Veröffentlicht: (2025)
Distributed Asynchronous Device Speech Enhancement via Windowed Cross-Attention
von: Yang, Gene-Ping, et al.
Veröffentlicht: (2025)
von: Yang, Gene-Ping, et al.
Veröffentlicht: (2025)
VoiceRestore: Flow-Matching Transformers for Speech Recording Quality Restoration
von: Kirdey, Stanislav
Veröffentlicht: (2025)
von: Kirdey, Stanislav
Veröffentlicht: (2025)
MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
Ultra-Low Latency Speech Enhancement - A Comprehensive Study
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
FLOWER: Flow-Based Estimated Gaussian Guidance for General Speech Restoration
von: Yang, Da-Hee, et al.
Veröffentlicht: (2025)
von: Yang, Da-Hee, et al.
Veröffentlicht: (2025)
MeanFlow-TSE: One-Step Generative Target Speaker Extraction with Mean Flow
von: Shimizu, Riki, et al.
Veröffentlicht: (2025)
von: Shimizu, Riki, et al.
Veröffentlicht: (2025)
Multi-label audio classification with a noisy zero-shot teacher
von: Braun, Sebastian, et al.
Veröffentlicht: (2024)
von: Braun, Sebastian, et al.
Veröffentlicht: (2024)
Gaussian Flow Bridges for Audio Domain Transfer with Unpaired Data
von: Moliner, Eloi, et al.
Veröffentlicht: (2024)
von: Moliner, Eloi, et al.
Veröffentlicht: (2024)
Neural Speech Coding for Real-time Communications using Constant Bitrate Scalar Quantization
von: Brendel, Andreas, et al.
Veröffentlicht: (2024)
von: Brendel, Andreas, et al.
Veröffentlicht: (2024)
Query-Based Asymmetric Modeling with Decoupled Input-Output Rates for Speech Restoration
von: Shin, Ui-Hyeop, et al.
Veröffentlicht: (2025)
von: Shin, Ui-Hyeop, et al.
Veröffentlicht: (2025)
Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
von: Byun, Kyungguen, et al.
Veröffentlicht: (2025)
von: Byun, Kyungguen, et al.
Veröffentlicht: (2025)
Real-time Stereo Speech Enhancement with Spatial-Cue Preservation based on Dual-Path Structure
von: Togami, Masahito, et al.
Veröffentlicht: (2024)
von: Togami, Masahito, et al.
Veröffentlicht: (2024)
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
von: Ku, Pin-Jui, et al.
Veröffentlicht: (2024)
A Fast Solver for Interpolating Stochastic Differential Equation Diffusion Models for Speech Restoration
von: Lay, Bunlong, et al.
Veröffentlicht: (2026)
von: Lay, Bunlong, et al.
Veröffentlicht: (2026)
Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of Speech
von: Liang, Xinyu, et al.
Veröffentlicht: (2025)
von: Liang, Xinyu, et al.
Veröffentlicht: (2025)
SALAD-VAE: Semantic Audio Compression with Language-Audio Distillation
von: Braun, Sebastian, et al.
Veröffentlicht: (2025)
von: Braun, Sebastian, et al.
Veröffentlicht: (2025)
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
Rethinking Flow and Diffusion Bridge Models for Speech Enhancement
von: Wang, Dahan, et al.
Veröffentlicht: (2026)
von: Wang, Dahan, et al.
Veröffentlicht: (2026)
The CCF AATC 2025 Speech Restoration Challenge: A Retrospective
von: Zhang, Junan, et al.
Veröffentlicht: (2025)
von: Zhang, Junan, et al.
Veröffentlicht: (2025)
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
FlowSE: Flow Matching-based Speech Enhancement
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
von: Lee, Seonggyu, et al.
Veröffentlicht: (2025)
Estimation and Restoration of Unknown Nonlinear Distortion using Diffusion
von: Švento, Michal, et al.
Veröffentlicht: (2025)
von: Švento, Michal, et al.
Veröffentlicht: (2025)
FlowSE-GRPO: Training Flow Matching Speech Enhancement via Online Reinforcement Learning
von: Wang, Haoxu, et al.
Veröffentlicht: (2026)
von: Wang, Haoxu, et al.
Veröffentlicht: (2026)
Evaluation of an ITD-to-ILD Transformation as a Method to Restore the Spatial Benefit in Speech Intelligibility in Hearing Impaired Listeners
von: Bäumer, Timm-Jonas, et al.
Veröffentlicht: (2025)
von: Bäumer, Timm-Jonas, et al.
Veröffentlicht: (2025)
Perceptual Ratings Predict Speech Inversion Articulatory Kinematics in Childhood Speech Sound Disorders
von: Benway, Nina R., et al.
Veröffentlicht: (2025)
von: Benway, Nina R., et al.
Veröffentlicht: (2025)
VC-ENHANCE: Speech Restoration with Integrated Noise Suppression and Voice Conversion
von: Byun, Kyungguen, et al.
Veröffentlicht: (2024)
von: Byun, Kyungguen, et al.
Veröffentlicht: (2024)
ReverbMiipher: Generative Speech Restoration meets Reverberation Characteristics Controllability
von: Nakata, Wataru, et al.
Veröffentlicht: (2025)
von: Nakata, Wataru, et al.
Veröffentlicht: (2025)
Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
Data Augmentation for Pathological Speech Enhancement
von: Hou, Mingchi, et al.
Veröffentlicht: (2026)
von: Hou, Mingchi, et al.
Veröffentlicht: (2026)
Adapting Frechet Audio Distance for Generative Music Evaluation
von: Gui, Azalea, et al.
Veröffentlicht: (2023)
von: Gui, Azalea, et al.
Veröffentlicht: (2023)
Miipher-2: A Universal Speech Restoration Model for Million-Hour Scale Data Restoration
von: Karita, Shigeki, et al.
Veröffentlicht: (2025)
von: Karita, Shigeki, et al.
Veröffentlicht: (2025)
Towards Frame-level Quality Predictions of Synthetic Speech
von: Kuhlmann, Michael, et al.
Veröffentlicht: (2025)
von: Kuhlmann, Michael, et al.
Veröffentlicht: (2025)
Restorative Speech Enhancement: A Progressive Approach Using SE and Codec Modules
von: Chiang, Hsin-Tien, et al.
Veröffentlicht: (2024)
von: Chiang, Hsin-Tien, et al.
Veröffentlicht: (2024)
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
von: Chung, Soo-Whan, et al.
Veröffentlicht: (2025)
von: Chung, Soo-Whan, et al.
Veröffentlicht: (2025)
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
von: Ma, Guobin, et al.
Veröffentlicht: (2025)
LP-CFM: Perceptual Invariance-Aware Conditional Flow Matching for Speech Modeling
von: Kwak, Doyeop, et al.
Veröffentlicht: (2025)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2025)
Modeling Multi-Level Hearing Loss for Speech Intelligibility Prediction
von: Zhou, Xiajie, et al.
Veröffentlicht: (2025)
von: Zhou, Xiajie, et al.
Veröffentlicht: (2025)
Generalizability of Predictive and Generative Speech Enhancement Models to Pathological Speakers
von: Hou, Mingchi, et al.
Veröffentlicht: (2025)
von: Hou, Mingchi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Real-Time Generative Speech Restoration with Flow-Matching
von: Hsieh, Tsun-An, et al.
Veröffentlicht: (2025) -
MeanSE: Efficient Generative Speech Enhancement with Mean Flows
von: Wang, Jiahe, et al.
Veröffentlicht: (2025) -
Distributed Asynchronous Device Speech Enhancement via Windowed Cross-Attention
von: Yang, Gene-Ping, et al.
Veröffentlicht: (2025) -
VoiceRestore: Flow-Matching Transformers for Speech Recording Quality Restoration
von: Kirdey, Stanislav
Veröffentlicht: (2025) -
MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
von: Zhu, Yike, et al.
Veröffentlicht: (2025)