PLDNet: PLD-Guided Lightweight Deep Network Boosted by Efficient Attention for Handheld Dual-Microphone Speech Enhancement
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Nan, Jiang, Youhai, Tan, Jialin, Qi, Chongmin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
di: Haeb-Umbach, Reinhold, et al.
Pubblicazione: (2025)
di: Haeb-Umbach, Reinhold, et al.
Pubblicazione: (2025)
HyBeam: Hybrid Microphone-Beamforming Array-Agnostic Speech Enhancement for Wearables
di: Ilan, Yuval Bar, et al.
Pubblicazione: (2025)
di: Ilan, Yuval Bar, et al.
Pubblicazione: (2025)
Impact of Microphone Array Mismatches to Learning-based Replay Speech Detection
di: Neri, Michael, et al.
Pubblicazione: (2025)
di: Neri, Michael, et al.
Pubblicazione: (2025)
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
di: Yan, Haoyin, et al.
Pubblicazione: (2024)
di: Yan, Haoyin, et al.
Pubblicazione: (2024)
Binaural Speech Enhancement Using Complex Convolutional Recurrent Networks
di: Tokala, Vikas, et al.
Pubblicazione: (2025)
di: Tokala, Vikas, et al.
Pubblicazione: (2025)
DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing
di: Lee, Dongheon, et al.
Pubblicazione: (2023)
di: Lee, Dongheon, et al.
Pubblicazione: (2023)
FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching
di: Wang, Ziqian, et al.
Pubblicazione: (2025)
di: Wang, Ziqian, et al.
Pubblicazione: (2025)
SELM: Speech Enhancement Using Discrete Tokens and Language Models
di: Wang, Ziqian, et al.
Pubblicazione: (2023)
di: Wang, Ziqian, et al.
Pubblicazione: (2023)
HOMULA-RIR: A Room Impulse Response Dataset for Teleconferencing and Spatial Audio Applications Acquired Through Higher-Order Microphones and Uniform Linear Microphone Arrays
di: Miotello, Federico, et al.
Pubblicazione: (2024)
di: Miotello, Federico, et al.
Pubblicazione: (2024)
Speech Enhancement based on cascaded two flows
di: Lee, Seonggyu, et al.
Pubblicazione: (2025)
di: Lee, Seonggyu, et al.
Pubblicazione: (2025)
FlowSE: Flow Matching-based Speech Enhancement
di: Lee, Seonggyu, et al.
Pubblicazione: (2025)
di: Lee, Seonggyu, et al.
Pubblicazione: (2025)
Unsupervised Face-Masked Speech Enhancement Using Generative Adversarial Networks With Human-in-the-Loop Assessment Metrics
di: Wang, Syu-Siang, et al.
Pubblicazione: (2024)
di: Wang, Syu-Siang, et al.
Pubblicazione: (2024)
Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
di: Bae, Hanbin, et al.
Pubblicazione: (2024)
di: Bae, Hanbin, et al.
Pubblicazione: (2024)
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
di: Kamo, Naoyuki, et al.
Pubblicazione: (2025)
di: Kamo, Naoyuki, et al.
Pubblicazione: (2025)
High-Density MIMO Localization Using a 32x64 Ultrasonic Transducer-Microphone Array with Real-Time Data Streaming
di: Baeyens, Rens, et al.
Pubblicazione: (2025)
di: Baeyens, Rens, et al.
Pubblicazione: (2025)
Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec
di: Ren, Yanzhou, et al.
Pubblicazione: (2026)
di: Ren, Yanzhou, et al.
Pubblicazione: (2026)
Steered Response Power-Based Direction-of-Arrival Estimation Exploiting an Auxiliary Microphone
di: Brümann, Klaus, et al.
Pubblicazione: (2024)
di: Brümann, Klaus, et al.
Pubblicazione: (2024)
BRUDEX Database: Binaural Room Impulse Responses with Uniformly Distributed External Microphones
di: Fejgin, Daniel, et al.
Pubblicazione: (2023)
di: Fejgin, Daniel, et al.
Pubblicazione: (2023)
Toward Universal Speech Enhancement for Diverse Input Conditions
di: Zhang, Wangyou, et al.
Pubblicazione: (2023)
di: Zhang, Wangyou, et al.
Pubblicazione: (2023)
SuperM2M: Supervised and Mixture-to-Mixture Co-Learning for Speech Enhancement and Noise-Robust ASR
di: Wang, Zhong-Qiu
Pubblicazione: (2024)
di: Wang, Zhong-Qiu
Pubblicazione: (2024)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
di: Wang, Kuan-Chen, et al.
Pubblicazione: (2024)
di: Wang, Kuan-Chen, et al.
Pubblicazione: (2024)
LABNet: A Lightweight Attentive Beamforming Network for Ad-hoc Multichannel Microphone Invariant Real-Time Speech Enhancement
di: Yan, Haoyin, et al.
Pubblicazione: (2025)
di: Yan, Haoyin, et al.
Pubblicazione: (2025)
Lessons Learned from the URGENT 2024 Speech Enhancement Challenge
di: Zhang, Wangyou, et al.
Pubblicazione: (2025)
di: Zhang, Wangyou, et al.
Pubblicazione: (2025)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
di: Sato, Hiroshi, et al.
Pubblicazione: (2025)
di: Sato, Hiroshi, et al.
Pubblicazione: (2025)
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
di: Serre, Thomas, et al.
Pubblicazione: (2026)
di: Serre, Thomas, et al.
Pubblicazione: (2026)
Exploiting an External Microphone for Binaural RTF-Vector-Based Direction of Arrival Estimation for Multiple Speakers
di: Fejgin, Daniel, et al.
Pubblicazione: (2023)
di: Fejgin, Daniel, et al.
Pubblicazione: (2023)
Parameter-Efficient Fine-Tuning of Foundation Models for CLP Speech Classification
di: Bhattacharjee, Susmita, et al.
Pubblicazione: (2025)
di: Bhattacharjee, Susmita, et al.
Pubblicazione: (2025)
Prompt-driven Target Speech Diarization
di: Jiang, Yidi, et al.
Pubblicazione: (2023)
di: Jiang, Yidi, et al.
Pubblicazione: (2023)
FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
di: Masuyama, Yoshiki, et al.
Pubblicazione: (2025)
di: Masuyama, Yoshiki, et al.
Pubblicazione: (2025)
HiRIS: an Airborne Sonar Sensor with a 1024 Channel Microphone Array for In-Air Acoustic Imaging
di: Laurijssen, Dennis, et al.
Pubblicazione: (2024)
di: Laurijssen, Dennis, et al.
Pubblicazione: (2024)
GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
di: Shetu, Shrishti Saha, et al.
Pubblicazione: (2024)
Advances in Microphone Array Processing and Multichannel Speech Enhancement
di: Huang, Gongping, et al.
Pubblicazione: (2025)
di: Huang, Gongping, et al.
Pubblicazione: (2025)
Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder
di: Xie, Yuying, et al.
Pubblicazione: (2024)
di: Xie, Yuying, et al.
Pubblicazione: (2024)
SpeechMLC: Speech Multi-label Classification
di: Kim, Miseul, et al.
Pubblicazione: (2025)
di: Kim, Miseul, et al.
Pubblicazione: (2025)
Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement
di: Yang, Yujie, et al.
Pubblicazione: (2025)
di: Yang, Yujie, et al.
Pubblicazione: (2025)
FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
di: Xu, Zhongweiyang, et al.
Pubblicazione: (2024)
di: Xu, Zhongweiyang, et al.
Pubblicazione: (2024)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
di: Qi, Tianhua, et al.
Pubblicazione: (2026)
di: Qi, Tianhua, et al.
Pubblicazione: (2026)
DeepFilterGAN: A Full-band Real-time Speech Enhancement System with GAN-based Stochastic Regeneration
di: Serbest, Sanberk, et al.
Pubblicazione: (2025)
di: Serbest, Sanberk, et al.
Pubblicazione: (2025)
A Speech Production Model for Radar: Connecting Speech Acoustics with Radar-Measured Vibrations
di: Lenz, Isabella, et al.
Pubblicazione: (2025)
di: Lenz, Isabella, et al.
Pubblicazione: (2025)
ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
di: Yang, Shu-wen, et al.
Pubblicazione: (2025)
di: Yang, Shu-wen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
di: Haeb-Umbach, Reinhold, et al.
Pubblicazione: (2025) -
HyBeam: Hybrid Microphone-Beamforming Array-Agnostic Speech Enhancement for Wearables
di: Ilan, Yuval Bar, et al.
Pubblicazione: (2025) -
Impact of Microphone Array Mismatches to Learning-based Replay Speech Detection
di: Neri, Michael, et al.
Pubblicazione: (2025) -
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
di: Yan, Haoyin, et al.
Pubblicazione: (2024) -
Binaural Speech Enhancement Using Complex Convolutional Recurrent Networks
di: Tokala, Vikas, et al.
Pubblicazione: (2025)