Real-time Stereo Speech Enhancement with Spatial-Cue Preservation based on Dual-Path Structure
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Togami, Masahito, Valin, Jean-Marc, Helwani, Karim, Giri, Ritwik, Isik, Umut, Goodwin, Michael M. |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
NoLACE: Improving Low-Complexity Speech Codec Enhancement Through Adaptive Temporal Shaping
par: Büthe, Jan, et autres
Publié: (2023)
par: Büthe, Jan, et autres
Publié: (2023)
Sound Source Separation Using Latent Variational Block-Wise Disentanglement
par: Helwani, Karim, et autres
Publié: (2024)
par: Helwani, Karim, et autres
Publié: (2024)
A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
par: Wang, Jingyuan, et autres
Publié: (2024)
par: Wang, Jingyuan, et autres
Publié: (2024)
RADE: A Neural Codec for Transmitting Speech over HF Radio Channels
par: Rowe, David, et autres
Publié: (2025)
par: Rowe, David, et autres
Publié: (2025)
A Lightweight Fourier-based Network for Binaural Speech Enhancement with Spatial Cue Preservation
par: Lu, Xikun, et autres
Publié: (2025)
par: Lu, Xikun, et autres
Publié: (2025)
Very Low Complexity Speech Synthesis Using Framewise Autoregressive GAN (FARGAN) with Pitch Prediction
par: Valin, Jean-Marc, et autres
Publié: (2024)
par: Valin, Jean-Marc, et autres
Publié: (2024)
DRED: Deep REDundancy Coding of Speech Using a Rate-Distortion-Optimized Variational Autoencoder
par: Valin, Jean-Marc, et autres
Publié: (2022)
par: Valin, Jean-Marc, et autres
Publié: (2022)
Noise-Robust DSP-Assisted Neural Pitch Estimation with Very Low Complexity
par: Subramani, Krishna, et autres
Publié: (2023)
par: Subramani, Krishna, et autres
Publié: (2023)
A lightweight and robust method for blind wideband-to-fullband extension of speech
par: Büthe, Jan, et autres
Publié: (2024)
par: Büthe, Jan, et autres
Publié: (2024)
WTFormer: A Wavelet Conformer Network for MIMO Speech Enhancement with Spatial Cues Peservation
par: Han, Lu, et autres
Publié: (2025)
par: Han, Lu, et autres
Publié: (2025)
ZipEnhancer: Dual-Path Down-Up Sampling-based Zipformer for Monaural Speech Enhancement
par: Wang, Haoxu, et autres
Publié: (2025)
par: Wang, Haoxu, et autres
Publié: (2025)
DAT-CFTNet: Speech Enhancement for Cochlear Implant Recipients using Attention-based Dual-Path Recurrent Neural Network
par: Mamun, Nursadul, et autres
Publié: (2026)
par: Mamun, Nursadul, et autres
Publié: (2026)
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
par: Yan, Haoyin, et autres
Publié: (2024)
par: Yan, Haoyin, et autres
Publié: (2024)
Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues
par: Hussain, Tassadaq, et autres
Publié: (2024)
par: Hussain, Tassadaq, et autres
Publié: (2024)
Audio-Visual Speech Enhancement for Spatial Audio - Spatial-VisualVoice and the MAVE Database
par: Yaffe, Danielle, et autres
Publié: (2025)
par: Yaffe, Danielle, et autres
Publié: (2025)
Direction-Preserving MIMO Speech Enhancement Using a Neural Covariance Estimator
par: Deppisch, Thomas
Publié: (2026)
par: Deppisch, Thomas
Publié: (2026)
Spatial-Filter-Bank-Based Neural Method for Multichannel Speech Enhancement
par: Zheng, Tianqin, et autres
Publié: (2025)
par: Zheng, Tianqin, et autres
Publié: (2025)
Universal Score-based Speech Enhancement with High Content Preservation
par: Scheibler, Robin, et autres
Publié: (2024)
par: Scheibler, Robin, et autres
Publié: (2024)
An Explicit Consistency-Preserving Loss Function for Phase Reconstruction and Speech Enhancement
par: Ku, Pin-Jui, et autres
Publié: (2024)
par: Ku, Pin-Jui, et autres
Publié: (2024)
Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative Conditions
par: La Quatra, Moreno, et autres
Publié: (2024)
par: La Quatra, Moreno, et autres
Publié: (2024)
A Variance-Preserving Interpolation Approach for Diffusion Models with Applications to Single Channel Speech Enhancement and Recognition
par: Guo, Zilu, et autres
Publié: (2024)
par: Guo, Zilu, et autres
Publié: (2024)
Decoupled Spatial and Temporal Processing for Resource Efficient Multichannel Speech Enhancement
par: Pandey, Ashutosh, et autres
Publié: (2024)
par: Pandey, Ashutosh, et autres
Publié: (2024)
Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
par: Ren, Wenze, et autres
Publié: (2024)
par: Ren, Wenze, et autres
Publié: (2024)
Leveraging Spatial Cues from Cochlear Implant Microphones to Efficiently Enhance Speech Separation in Real-World Listening Scenes
par: Olalere, Feyisayo, et autres
Publié: (2025)
par: Olalere, Feyisayo, et autres
Publié: (2025)
A Lightweight Hybrid Dual Channel Speech Enhancement System under Low-SNR Conditions
par: Wang, Zheng, et autres
Publié: (2025)
par: Wang, Zheng, et autres
Publié: (2025)
Influence of Clean Speech Characteristics on Speech Enhancement Performance
par: Hou, Mingchi, et autres
Publié: (2025)
par: Hou, Mingchi, et autres
Publié: (2025)
DroFiT: A Lightweight Band-fused Frequency Attention Toward Real-time UAV Speech Enhancement
par: Lee, Jeongmin, et autres
Publié: (2025)
par: Lee, Jeongmin, et autres
Publié: (2025)
A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition
par: de Groot, Dimme, et autres
Publié: (2026)
par: de Groot, Dimme, et autres
Publié: (2026)
Inter-Speaker Relative Cues for Two-Stage Text-Guided Target Speech Extraction
par: Dai, Wang, et autres
Publié: (2026)
par: Dai, Wang, et autres
Publié: (2026)
Assessing the Impact of Noise and Speech Enhancement on the Intelligibility of Speech Codecs
par: Behringer, Lyonel, et autres
Publié: (2026)
par: Behringer, Lyonel, et autres
Publié: (2026)
Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario
par: Wen, Wen, et autres
Publié: (2024)
par: Wen, Wen, et autres
Publié: (2024)
Conditional Latent Diffusion-Based Speech Enhancement Via Dual Context Learning
par: Zhao, Shengkui, et autres
Publié: (2025)
par: Zhao, Shengkui, et autres
Publié: (2025)
Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
par: Khan, Muhammad Salman, et autres
Publié: (2024)
par: Khan, Muhammad Salman, et autres
Publié: (2024)
Exploring Efficient Directional and Distance Cues for Regional Speech Separation
par: Jiang, Yiheng, et autres
Publié: (2025)
par: Jiang, Yiheng, et autres
Publié: (2025)
Hybrid Real- And Complex-Valued Neural Network Concept For Low-Complexity Phase-Aware Speech Enhancement
par: Fiorio, Luan Vinícius, et autres
Publié: (2025)
par: Fiorio, Luan Vinícius, et autres
Publié: (2025)
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
par: Ochiai, Tsubasa, et autres
Publié: (2024)
par: Ochiai, Tsubasa, et autres
Publié: (2024)
Data Augmentation for Pathological Speech Enhancement
par: Hou, Mingchi, et autres
Publié: (2026)
par: Hou, Mingchi, et autres
Publié: (2026)
Schrödinger Bridge for Generative Speech Enhancement
par: Jukić, Ante, et autres
Publié: (2024)
par: Jukić, Ante, et autres
Publié: (2024)
RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization
par: Yang, Bing, et autres
Publié: (2024)
par: Yang, Bing, et autres
Publié: (2024)
A Dual-Branch Parallel Network for Speech Enhancement and Restoration
par: Yang, Da-Hee, et autres
Publié: (2024)
par: Yang, Da-Hee, et autres
Publié: (2024)
Documents similaires
-
NoLACE: Improving Low-Complexity Speech Codec Enhancement Through Adaptive Temporal Shaping
par: Büthe, Jan, et autres
Publié: (2023) -
Sound Source Separation Using Latent Variational Block-Wise Disentanglement
par: Helwani, Karim, et autres
Publié: (2024) -
A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
par: Wang, Jingyuan, et autres
Publié: (2024) -
RADE: A Neural Codec for Transmitting Speech over HF Radio Channels
par: Rowe, David, et autres
Publié: (2025) -
A Lightweight Fourier-based Network for Binaural Speech Enhancement with Spatial Cue Preservation
par: Lu, Xikun, et autres
Publié: (2025)