What do neural networks listen to? Exploring the crucial bands in Speech Enhancement using Sinc-convolution
Fuente:
arXiv
Guardado en:
| Autores principales: | Ho, Kuan-Hsun, Hung, Jeih-weih, Chen, Berlin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ConSep: a Noise- and Reverberation-Robust Speech Separation Framework by Magnitude Conditioning
por: Ho, Kuan-Hsun, et al.
Publicado: (2024)
por: Ho, Kuan-Hsun, et al.
Publicado: (2024)
Multispecies bird sound recognition using a fully convolutional neural network
por: García-Ordás, María Teresa, et al.
Publicado: (2024)
por: García-Ordás, María Teresa, et al.
Publicado: (2024)
Towards detecting the pathological subharmonic voicing with fully convolutional neural networks
por: Ikuma, Takeshi, et al.
Publicado: (2025)
por: Ikuma, Takeshi, et al.
Publicado: (2025)
Multimodal Emotion Recognition from Raw Audio with Sinc-convolution
por: Zhang, Xiaohui, et al.
Publicado: (2024)
por: Zhang, Xiaohui, et al.
Publicado: (2024)
Universal Robust Speech Adaptation for Cross-Domain Speech Recognition and Enhancement
por: Wang, Chien-Chun, et al.
Publicado: (2026)
por: Wang, Chien-Chun, et al.
Publicado: (2026)
Speech privacy-preserving methods using secret key for convolutional neural network models and their robustness evaluation
por: Niwa, Shoko, et al.
Publicado: (2024)
por: Niwa, Shoko, et al.
Publicado: (2024)
Speech-Aware Neural Diarization with Encoder-Decoder Attractor Guided by Attention Constraints
por: Lee, PeiYing, et al.
Publicado: (2024)
por: Lee, PeiYing, et al.
Publicado: (2024)
Improving Speech Enhancement by Cross- and Sub-band Processing with State Space Model
por: Li, Jizhen, et al.
Publicado: (2025)
por: Li, Jizhen, et al.
Publicado: (2025)
Adaptive high-precision sound source localization at low frequencies based on convolutional neural network
por: Ma, Wenbo, et al.
Publicado: (2024)
por: Ma, Wenbo, et al.
Publicado: (2024)
Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI
por: Lin, Yi-Cheng, et al.
Publicado: (2026)
por: Lin, Yi-Cheng, et al.
Publicado: (2026)
Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network
por: Chen, Yanan, et al.
Publicado: (2024)
por: Chen, Yanan, et al.
Publicado: (2024)
Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
por: Ren, Wenze, et al.
Publicado: (2024)
por: Ren, Wenze, et al.
Publicado: (2024)
GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
por: Yao, Jixun, et al.
Publicado: (2025)
por: Yao, Jixun, et al.
Publicado: (2025)
An Explicit Consistency-Preserving Loss Function for Phase Reconstruction and Speech Enhancement
por: Ku, Pin-Jui, et al.
Publicado: (2024)
por: Ku, Pin-Jui, et al.
Publicado: (2024)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
por: Wang, Kuan-Chen, et al.
Publicado: (2024)
por: Wang, Kuan-Chen, et al.
Publicado: (2024)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
por: Chao, Rong, et al.
Publicado: (2025)
por: Chao, Rong, et al.
Publicado: (2025)
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
por: Ren, Wenze, et al.
Publicado: (2024)
por: Ren, Wenze, et al.
Publicado: (2024)
Attention-Based Beamformer For Multi-Channel Speech Enhancement
por: Bai, Jinglin, et al.
Publicado: (2024)
por: Bai, Jinglin, et al.
Publicado: (2024)
Dataset-Distillation Generative Model for Speech Emotion Recognition
por: Ritter-Gutierrez, Fabian, et al.
Publicado: (2024)
por: Ritter-Gutierrez, Fabian, et al.
Publicado: (2024)
Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
por: Khan, Muhammad Salman, et al.
Publicado: (2024)
por: Khan, Muhammad Salman, et al.
Publicado: (2024)
Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement
por: Nasretdinov, Rauf, et al.
Publicado: (2025)
por: Nasretdinov, Rauf, et al.
Publicado: (2025)
Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement
por: Yuan, Xihao, et al.
Publicado: (2025)
por: Yuan, Xihao, et al.
Publicado: (2025)
SaD: A Scenario-Aware Discriminator for Speech Enhancement
por: Yuan, Xihao, et al.
Publicado: (2025)
por: Yuan, Xihao, et al.
Publicado: (2025)
FRCRN: Boosting Feature Representation using Frequency Recurrence for Monaural Speech Enhancement
por: Zhao, Shengkui, et al.
Publicado: (2022)
por: Zhao, Shengkui, et al.
Publicado: (2022)
Absorbing Discrete Diffusion for Speech Enhancement
por: Gonzalez, Philippe
Publicado: (2026)
por: Gonzalez, Philippe
Publicado: (2026)
Effective Noise-aware Data Simulation for Domain-adaptive Speech Enhancement Leveraging Dynamic Stochastic Perturbation
por: Wang, Chien-Chun, et al.
Publicado: (2024)
por: Wang, Chien-Chun, et al.
Publicado: (2024)
Toward end-to-end interpretable convolutional neural networks for waveform signals
por: Vu, Linh, et al.
Publicado: (2024)
por: Vu, Linh, et al.
Publicado: (2024)
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
por: Ma, Ding, et al.
Publicado: (2026)
por: Ma, Ding, et al.
Publicado: (2026)
Objective and Subjective Evaluation of Diffusion-Based Speech Enhancement for Dysarthric Speech
por: de Groot, Dimme, et al.
Publicado: (2025)
por: de Groot, Dimme, et al.
Publicado: (2025)
Speech foundation models on intelligibility prediction for hearing-impaired listeners
por: Cuervo, Santiago, et al.
Publicado: (2024)
por: Cuervo, Santiago, et al.
Publicado: (2024)
ICASSP 2026 URGENT Speech Enhancement Challenge
por: Li, Chenda, et al.
Publicado: (2026)
por: Li, Chenda, et al.
Publicado: (2026)
Investigating Training Objectives for Generative Speech Enhancement
por: Richter, Julius, et al.
Publicado: (2024)
por: Richter, Julius, et al.
Publicado: (2024)
DISPATCH: Distilling Selective Patches for Speech Enhancement
por: Kim, Dohwan, et al.
Publicado: (2025)
por: Kim, Dohwan, et al.
Publicado: (2025)
Geneses: Unified Generative Speech Enhancement and Separation
por: Asai, Kohei, et al.
Publicado: (2026)
por: Asai, Kohei, et al.
Publicado: (2026)
Variational Autoencoder for Personalized Pathological Speech Enhancement
por: Hou, Mingchi, et al.
Publicado: (2025)
por: Hou, Mingchi, et al.
Publicado: (2025)
Universal Speech Enhancement with Regression and Generative Mamba
por: Chao, Rong, et al.
Publicado: (2025)
por: Chao, Rong, et al.
Publicado: (2025)
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
por: Yan, Haoyin, et al.
Publicado: (2024)
por: Yan, Haoyin, et al.
Publicado: (2024)
Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
por: Sutherland, Robert, et al.
Publicado: (2024)
por: Sutherland, Robert, et al.
Publicado: (2024)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
por: Huang, Kuan-Po, et al.
Publicado: (2023)
por: Huang, Kuan-Po, et al.
Publicado: (2023)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
por: Han, Seungu, et al.
Publicado: (2026)
por: Han, Seungu, et al.
Publicado: (2026)
Ejemplares similares
-
ConSep: a Noise- and Reverberation-Robust Speech Separation Framework by Magnitude Conditioning
por: Ho, Kuan-Hsun, et al.
Publicado: (2024) -
Multispecies bird sound recognition using a fully convolutional neural network
por: García-Ordás, María Teresa, et al.
Publicado: (2024) -
Towards detecting the pathological subharmonic voicing with fully convolutional neural networks
por: Ikuma, Takeshi, et al.
Publicado: (2025) -
Multimodal Emotion Recognition from Raw Audio with Sinc-convolution
por: Zhang, Xiaohui, et al.
Publicado: (2024) -
Universal Robust Speech Adaptation for Cross-Domain Speech Recognition and Enhancement
por: Wang, Chien-Chun, et al.
Publicado: (2026)