Improving Speech Enhancement by Cross- and Sub-band Processing with State Space Model
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Jizhen, Tu, Weiping, Yang, Yuhong, Xu, Xinmeng, Zhang, Yiqun, Ren, Yanzhen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Improving Speech Enhancement by Integrating Inter-Channel and Band Features with Dual-branch Conformer
di: Li, Jizhen, et al.
Pubblicazione: (2024)
di: Li, Jizhen, et al.
Pubblicazione: (2024)
SE Territory: Monaural Speech Enhancement Meets the Fixed Virtual Perceptual Space Mapping
di: Xu, Xinmeng, et al.
Pubblicazione: (2023)
di: Xu, Xinmeng, et al.
Pubblicazione: (2023)
SuperCodec: A Neural Speech Codec with Selective Back-Projection Network
di: Zheng, Youqiang, et al.
Pubblicazione: (2024)
di: Zheng, Youqiang, et al.
Pubblicazione: (2024)
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
di: Yan, Haoyin, et al.
Pubblicazione: (2024)
di: Yan, Haoyin, et al.
Pubblicazione: (2024)
SICRN: Advancing Speech Enhancement through State Space Model and Inplace Convolution Techniques
di: Zhao, Changjiang, et al.
Pubblicazione: (2024)
di: Zhao, Changjiang, et al.
Pubblicazione: (2024)
EMALG: An Enhanced Mandarin Lombard Grid Corpus with Meaningful Sentences
di: Li, Baifeng, et al.
Pubblicazione: (2023)
di: Li, Baifeng, et al.
Pubblicazione: (2023)
Exploring Sentence Type Effects on the Lombard Effect and Intelligibility Enhancement: A Comparative Study of Natural and Grid Sentences
di: Chen, Hongyang, et al.
Pubblicazione: (2023)
di: Chen, Hongyang, et al.
Pubblicazione: (2023)
Semantic Proximity Alignment: Towards Human Perception-consistent Audio Tagging by Aligning with Label Text Description
di: Liu, Wuyang, et al.
Pubblicazione: (2023)
di: Liu, Wuyang, et al.
Pubblicazione: (2023)
Towards Ultra-Low-Power Neuromorphic Speech Enhancement with Spiking-FullSubNet
di: Hao, Xiang, et al.
Pubblicazione: (2024)
di: Hao, Xiang, et al.
Pubblicazione: (2024)
Decoupled Spatial and Temporal Processing for Resource Efficient Multichannel Speech Enhancement
di: Pandey, Ashutosh, et al.
Pubblicazione: (2024)
di: Pandey, Ashutosh, et al.
Pubblicazione: (2024)
Improving Design of Input Condition Invariant Speech Enhancement
di: Zhang, Wangyou, et al.
Pubblicazione: (2024)
di: Zhang, Wangyou, et al.
Pubblicazione: (2024)
Sub-band and Full-band Interactive U-Net with DPRNN for Demixing Cross-talk Stereo Music
di: Yin, Han, et al.
Pubblicazione: (2024)
di: Yin, Han, et al.
Pubblicazione: (2024)
EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection
di: Zhang, Tong, et al.
Pubblicazione: (2025)
di: Zhang, Tong, et al.
Pubblicazione: (2025)
FreeCodec: A disentangled neural speech codec with fewer tokens
di: Zheng, Youqiang, et al.
Pubblicazione: (2024)
di: Zheng, Youqiang, et al.
Pubblicazione: (2024)
SimuSOE: A Simulated Snoring Dataset for Obstructive Sleep Apnea-Hypopnea Syndrome Evaluation during Wakefulness
di: Lin, Jie, et al.
Pubblicazione: (2024)
di: Lin, Jie, et al.
Pubblicazione: (2024)
From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models
di: Mu, Zhaoxi, et al.
Pubblicazione: (2025)
di: Mu, Zhaoxi, et al.
Pubblicazione: (2025)
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
di: Ren, Wenze, et al.
Pubblicazione: (2024)
di: Ren, Wenze, et al.
Pubblicazione: (2024)
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
Improved Remixing Process for Domain Adaptation-Based Speech Enhancement by Mitigating Data Imbalance in Signal-to-Noise Ratio
di: Li, Li, et al.
Pubblicazione: (2024)
di: Li, Li, et al.
Pubblicazione: (2024)
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
di: Ochiai, Tsubasa, et al.
Pubblicazione: (2024)
di: Ochiai, Tsubasa, et al.
Pubblicazione: (2024)
BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech Enhancement
di: Fan, Cunhang, et al.
Pubblicazione: (2024)
di: Fan, Cunhang, et al.
Pubblicazione: (2024)
A Two-Stage Framework in Cross-Spectrum Domain for Real-Time Speech Enhancement
di: Zhang, Yuewei, et al.
Pubblicazione: (2024)
di: Zhang, Yuewei, et al.
Pubblicazione: (2024)
Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment
di: Wang, Wei, et al.
Pubblicazione: (2025)
di: Wang, Wei, et al.
Pubblicazione: (2025)
What do neural networks listen to? Exploring the crucial bands in Speech Enhancement using Sinc-convolution
di: Ho, Kuan-Hsun, et al.
Pubblicazione: (2024)
di: Ho, Kuan-Hsun, et al.
Pubblicazione: (2024)
Improving Resource-Efficient Speech Enhancement via Neural Differentiable DSP Vocoder Refinement
di: Guimarães, Heitor R., et al.
Pubblicazione: (2025)
di: Guimarães, Heitor R., et al.
Pubblicazione: (2025)
Complex-Cycle-Consistent Diffusion Model for Monaural Speech Enhancement
di: Li, Yi, et al.
Pubblicazione: (2024)
di: Li, Yi, et al.
Pubblicazione: (2024)
Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network
di: Chen, Yanan, et al.
Pubblicazione: (2024)
di: Chen, Yanan, et al.
Pubblicazione: (2024)
State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition
di: Farhadipour, Aref, et al.
Pubblicazione: (2025)
di: Farhadipour, Aref, et al.
Pubblicazione: (2025)
FullSubNet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement
di: Hao, Xiang, et al.
Pubblicazione: (2020)
di: Hao, Xiang, et al.
Pubblicazione: (2020)
Geometry-Constrained EEG Channel Selection for Brain-Assisted Speech Enhancement
di: Zuo, Keying, et al.
Pubblicazione: (2024)
di: Zuo, Keying, et al.
Pubblicazione: (2024)
Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
di: Sutherland, Robert, et al.
Pubblicazione: (2024)
di: Sutherland, Robert, et al.
Pubblicazione: (2024)
Multi-modal Speech Enhancement with Limited Electromyography Channels
di: Feng, Fuyuan, et al.
Pubblicazione: (2025)
di: Feng, Fuyuan, et al.
Pubblicazione: (2025)
Adaptive Convolution for CNN-based Speech Enhancement Models
di: Wang, Dahan, et al.
Pubblicazione: (2025)
di: Wang, Dahan, et al.
Pubblicazione: (2025)
Schrödinger Bridge Consistency Trajectory Models for Speech Enhancement
di: Nishigori, Shuichiro, et al.
Pubblicazione: (2025)
di: Nishigori, Shuichiro, et al.
Pubblicazione: (2025)
A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model
di: Xiang, Yang, et al.
Pubblicazione: (2025)
di: Xiang, Yang, et al.
Pubblicazione: (2025)
Robust One-step Speech Enhancement via Consistency Distillation
di: Xu, Liang, et al.
Pubblicazione: (2025)
di: Xu, Liang, et al.
Pubblicazione: (2025)
Restorative Speech Enhancement: A Progressive Approach Using SE and Codec Modules
di: Chiang, Hsin-Tien, et al.
Pubblicazione: (2024)
di: Chiang, Hsin-Tien, et al.
Pubblicazione: (2024)
Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model
di: Ueda, Lucas, et al.
Pubblicazione: (2025)
di: Ueda, Lucas, et al.
Pubblicazione: (2025)
Spiking Structured State Space Model for Monaural Speech Enhancement
di: Du, Yu, et al.
Pubblicazione: (2023)
di: Du, Yu, et al.
Pubblicazione: (2023)
NoLACE: Improving Low-Complexity Speech Codec Enhancement Through Adaptive Temporal Shaping
di: Büthe, Jan, et al.
Pubblicazione: (2023)
di: Büthe, Jan, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Improving Speech Enhancement by Integrating Inter-Channel and Band Features with Dual-branch Conformer
di: Li, Jizhen, et al.
Pubblicazione: (2024) -
SE Territory: Monaural Speech Enhancement Meets the Fixed Virtual Perceptual Space Mapping
di: Xu, Xinmeng, et al.
Pubblicazione: (2023) -
SuperCodec: A Neural Speech Codec with Selective Back-Projection Network
di: Zheng, Youqiang, et al.
Pubblicazione: (2024) -
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
di: Yan, Haoyin, et al.
Pubblicazione: (2024) -
SICRN: Advancing Speech Enhancement through State Space Model and Inplace Convolution Techniques
di: Zhao, Changjiang, et al.
Pubblicazione: (2024)