RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Mingshuai, Chen, Zhuangqi, Yan, Xiaopeng, Lv, Yuanjun, Xia, Xianjun, Huang, Chuanzeng, Xiao, Yijian, Xie, Lei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RaD-Net: A Repairing and Denoising Network for Speech Signal Improvement
di: Liu, Mingshuai, et al.
Pubblicazione: (2024)
di: Liu, Mingshuai, et al.
Pubblicazione: (2024)
BS-PLCNet 2: Two-stage Band-split Packet Loss Concealment Network with Intra-model Knowledge Distillation
di: Zhang, Zihan, et al.
Pubblicazione: (2024)
di: Zhang, Zihan, et al.
Pubblicazione: (2024)
BS-PLCNet: Band-split Packet Loss Concealment Network with Multi-task Learning Framework and Multi-discriminators
di: Zhang, Zihan, et al.
Pubblicazione: (2024)
di: Zhang, Zihan, et al.
Pubblicazione: (2024)
S$^2$Voice: Style-Aware Autoregressive Modeling with Enhanced Conditioning for Singing Style Conversion
di: Wang, Ziqian, et al.
Pubblicazione: (2026)
di: Wang, Ziqian, et al.
Pubblicazione: (2026)
LDCodec: A high quality neural audio codec with low-complexity decoder
di: Jiang, Jiawei, et al.
Pubblicazione: (2025)
di: Jiang, Jiawei, et al.
Pubblicazione: (2025)
A lightweight dual-stage framework for personalized speech enhancement based on DeepFilterNet2
di: Serre, Thomas, et al.
Pubblicazione: (2024)
di: Serre, Thomas, et al.
Pubblicazione: (2024)
Triage knowledge distillation for speaker verification
di: Kim, Ju-ho, et al.
Pubblicazione: (2026)
di: Kim, Ju-ho, et al.
Pubblicazione: (2026)
Inter-channel Conv-TasNet for multichannel speech enhancement
di: Lee, Dongheon, et al.
Pubblicazione: (2021)
di: Lee, Dongheon, et al.
Pubblicazione: (2021)
UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling
di: Wang, Ziqian, et al.
Pubblicazione: (2025)
di: Wang, Ziqian, et al.
Pubblicazione: (2025)
Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios
di: Huang, Ziling, et al.
Pubblicazione: (2025)
di: Huang, Ziling, et al.
Pubblicazione: (2025)
Omni-directional attention mechanism based on Mamba for speech separation
di: Xue, Ke, et al.
Pubblicazione: (2026)
di: Xue, Ke, et al.
Pubblicazione: (2026)
Explainable speech emotion recognition through attentive pooling: insights from attention-based temporal localization
di: Leygue, Tahitoa, et al.
Pubblicazione: (2025)
di: Leygue, Tahitoa, et al.
Pubblicazione: (2025)
Comparison of linear and nonlinear methods for decoding selective attention to speech from ear-EEG recordings
di: Thornton, Mike, et al.
Pubblicazione: (2024)
di: Thornton, Mike, et al.
Pubblicazione: (2024)
Spatial-Magnifier: Spatial upsampling for multichannel speech enhancement
di: Lee, Dongheon, et al.
Pubblicazione: (2026)
di: Lee, Dongheon, et al.
Pubblicazione: (2026)
KS-Net: Multi-band joint speech restoration and enhancement network for 2024 ICASSP SSI Challenge
di: Yu, Guochen, et al.
Pubblicazione: (2024)
di: Yu, Guochen, et al.
Pubblicazione: (2024)
Predicting speech intelligibility in older adults for speech enhancement using the Gammachirp Envelope Similarity Index, GESI
di: Yamamoto, Ayako, et al.
Pubblicazione: (2025)
di: Yamamoto, Ayako, et al.
Pubblicazione: (2025)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
Real-time multichannel deep speech enhancement in hearing aids: Comparing monaural and binaural processing in complex acoustic scenarios
di: Westhausen, Nils L., et al.
Pubblicazione: (2024)
di: Westhausen, Nils L., et al.
Pubblicazione: (2024)
Fast-ULCNet: A fast and ultra low complexity network for single-channel speech enhancement
di: Larraza, Nicolás Arrieta, et al.
Pubblicazione: (2026)
di: Larraza, Nicolás Arrieta, et al.
Pubblicazione: (2026)
Heterogeneous bimodal attention fusion for speech emotion recognition
di: Luo, Jiachen, et al.
Pubblicazione: (2025)
di: Luo, Jiachen, et al.
Pubblicazione: (2025)
A low latency attention module for streaming self-supervised speech representation learning
di: Ma, Jianbo, et al.
Pubblicazione: (2023)
di: Ma, Jianbo, et al.
Pubblicazione: (2023)
Spatially constrained vs. unconstrained filtering in neural spatiospectral filters for multichannel speech enhancement
di: Briegleb, Annika, et al.
Pubblicazione: (2024)
di: Briegleb, Annika, et al.
Pubblicazione: (2024)
Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy
di: Ma, Linhan, et al.
Pubblicazione: (2024)
di: Ma, Linhan, et al.
Pubblicazione: (2024)
TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet
di: Jeong, Jaeseok, et al.
Pubblicazione: (2025)
di: Jeong, Jaeseok, et al.
Pubblicazione: (2025)
Unsupervised speech enhancement with spectral kurtosis and double deep priors
di: Ohnaka, Hien, et al.
Pubblicazione: (2024)
di: Ohnaka, Hien, et al.
Pubblicazione: (2024)
SPGM: Prioritizing Local Features for enhanced speech separation performance
di: Yip, Jia Qi, et al.
Pubblicazione: (2023)
di: Yip, Jia Qi, et al.
Pubblicazione: (2023)
On the social bias of speech self-supervised models
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
TokenSE: a Mamba-based discrete token speech enhancement framework for cochlear implants
di: Chiang, Hsin-Tien, et al.
Pubblicazione: (2026)
di: Chiang, Hsin-Tien, et al.
Pubblicazione: (2026)
Relational graph-driven differential denoising and diffusion attention fusion for multimodal conversation emotion recognition
di: Liu, Ying, et al.
Pubblicazione: (2026)
di: Liu, Ying, et al.
Pubblicazione: (2026)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
di: Pascu, Octavian, et al.
Pubblicazione: (2023)
di: Pascu, Octavian, et al.
Pubblicazione: (2023)
AlignNet: Learning dataset score alignment functions to enable better training of speech quality estimators
di: Pieper, Jaden, et al.
Pubblicazione: (2024)
di: Pieper, Jaden, et al.
Pubblicazione: (2024)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
di: Gowda, Harshavardhana T., et al.
Pubblicazione: (2025)
di: Gowda, Harshavardhana T., et al.
Pubblicazione: (2025)
GDiffuSE: Diffusion-based speech enhancement with noise model guidance
di: Yanir, Efrayim, et al.
Pubblicazione: (2025)
di: Yanir, Efrayim, et al.
Pubblicazione: (2025)
Monaural speech enhancement on drone via Adapter based transfer learning
di: Chen, Xingyu, et al.
Pubblicazione: (2024)
di: Chen, Xingyu, et al.
Pubblicazione: (2024)
DBMIF: a deep balanced multimodal iterative fusion framework for air- and bone-conduction speech enhancement
di: Wu, Yilei, et al.
Pubblicazione: (2026)
di: Wu, Yilei, et al.
Pubblicazione: (2026)
Real-time speech enhancement in noise for throat microphone using neural audio codec as foundation model
di: Hauret, Julien, et al.
Pubblicazione: (2025)
di: Hauret, Julien, et al.
Pubblicazione: (2025)
On combining acoustic and modulation spectrograms in an attention LSTM-based system for speech intelligibility level classification
di: Gallardo-Antolín, Ascensión, et al.
Pubblicazione: (2024)
di: Gallardo-Antolín, Ascensión, et al.
Pubblicazione: (2024)
Modeling strategies for speech enhancement in the latent space of a neural audio codec
di: Kammoun, Sofiene, et al.
Pubblicazione: (2025)
di: Kammoun, Sofiene, et al.
Pubblicazione: (2025)
Using RLHF to align speech enhancement approaches to mean-opinion quality scores
di: Kumar, Anurag, et al.
Pubblicazione: (2024)
di: Kumar, Anurag, et al.
Pubblicazione: (2024)
Single-channel speech enhancement by using psychoacoustical model inspired fusion framework
di: Samui, Suman
Pubblicazione: (2022)
di: Samui, Suman
Pubblicazione: (2022)
Documenti analoghi
-
RaD-Net: A Repairing and Denoising Network for Speech Signal Improvement
di: Liu, Mingshuai, et al.
Pubblicazione: (2024) -
BS-PLCNet 2: Two-stage Band-split Packet Loss Concealment Network with Intra-model Knowledge Distillation
di: Zhang, Zihan, et al.
Pubblicazione: (2024) -
BS-PLCNet: Band-split Packet Loss Concealment Network with Multi-task Learning Framework and Multi-discriminators
di: Zhang, Zihan, et al.
Pubblicazione: (2024) -
S$^2$Voice: Style-Aware Autoregressive Modeling with Enhanced Conditioning for Singing Style Conversion
di: Wang, Ziqian, et al.
Pubblicazione: (2026) -
LDCodec: A high quality neural audio codec with low-complexity decoder
di: Jiang, Jiawei, et al.
Pubblicazione: (2025)