Microphone Array Signal Processing and Deep Learning for Speech Enhancement
Fuente:
arXiv
Salvato in:
| Autori principali: | Haeb-Umbach, Reinhold, Nakatani, Tomohiro, Delcroix, Marc, Boeddeker, Christoph, Ochiai, Tsubasa |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization
di: von Neumann, Thilo, et al.
Pubblicazione: (2023)
di: von Neumann, Thilo, et al.
Pubblicazione: (2023)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
di: Sato, Hiroshi, et al.
Pubblicazione: (2025)
di: Sato, Hiroshi, et al.
Pubblicazione: (2025)
Array Geometry-Robust Attention-Based Neural Beamformer for Moving Speakers
di: Tammen, Marvin, et al.
Pubblicazione: (2024)
di: Tammen, Marvin, et al.
Pubblicazione: (2024)
Word Error Rate Definitions and Algorithms for Long-Form Multi-talker Speech Recognition
di: von Neumann, Thilo, et al.
Pubblicazione: (2025)
di: von Neumann, Thilo, et al.
Pubblicazione: (2025)
Loose coupling of spectral and spatial models for multi-channel diarization and enhancement of meetings in dynamic environments
di: Meise, Adrian, et al.
Pubblicazione: (2026)
di: Meise, Adrian, et al.
Pubblicazione: (2026)
30+ Years of Source Separation Research: Achievements and Future Challenges
di: Araki, Shoko, et al.
Pubblicazione: (2025)
di: Araki, Shoko, et al.
Pubblicazione: (2025)
TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings
di: Boeddeker, Christoph, et al.
Pubblicazione: (2023)
di: Boeddeker, Christoph, et al.
Pubblicazione: (2023)
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
di: Ochiai, Tsubasa, et al.
Pubblicazione: (2024)
di: Ochiai, Tsubasa, et al.
Pubblicazione: (2024)
Interaural time difference loss for binaural target sound extraction
di: Hernandez-Olivan, Carlos, et al.
Pubblicazione: (2024)
di: Hernandez-Olivan, Carlos, et al.
Pubblicazione: (2024)
MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems
di: von Neumann, Thilo, et al.
Pubblicazione: (2023)
di: von Neumann, Thilo, et al.
Pubblicazione: (2023)
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
di: Sato, Hiroshi, et al.
Pubblicazione: (2024)
di: Sato, Hiroshi, et al.
Pubblicazione: (2024)
Reference Microphone Selection for Guided Source Separation based on the Normalized L-p Norm
di: Lohmann, Anselm, et al.
Pubblicazione: (2025)
di: Lohmann, Anselm, et al.
Pubblicazione: (2025)
SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
di: Hernandez-Olivan, Carlos, et al.
Pubblicazione: (2024)
di: Hernandez-Olivan, Carlos, et al.
Pubblicazione: (2024)
Frontend Token Enhancement for Token-Based Speech Recognition
di: Ashihara, Takanori, et al.
Pubblicazione: (2026)
di: Ashihara, Takanori, et al.
Pubblicazione: (2026)
FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
di: Masuyama, Yoshiki, et al.
Pubblicazione: (2025)
di: Masuyama, Yoshiki, et al.
Pubblicazione: (2025)
Target Speech Extraction with Pre-trained Self-supervised Learning Models
di: Peng, Junyi, et al.
Pubblicazione: (2024)
di: Peng, Junyi, et al.
Pubblicazione: (2024)
Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2025)
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2025)
Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture Models
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2024)
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2024)
Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment
di: Boeddeker, Christoph, et al.
Pubblicazione: (2024)
di: Boeddeker, Christoph, et al.
Pubblicazione: (2024)
Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
di: Vieting, Peter, et al.
Pubblicazione: (2023)
di: Vieting, Peter, et al.
Pubblicazione: (2023)
Diminishing Domain Mismatch for DNN-Based Acoustic Distance Estimation via Stochastic Room Reverberation Models
di: Gburrek, Tobias, et al.
Pubblicazione: (2024)
di: Gburrek, Tobias, et al.
Pubblicazione: (2024)
Probing Self-supervised Learning Models with Target Speech Extraction
di: Peng, Junyi, et al.
Pubblicazione: (2024)
di: Peng, Junyi, et al.
Pubblicazione: (2024)
Speech Synthesis along Perceptual Voice Quality Dimensions
di: Rautenberg, Frederik, et al.
Pubblicazione: (2025)
di: Rautenberg, Frederik, et al.
Pubblicazione: (2025)
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
di: Kamo, Naoyuki, et al.
Pubblicazione: (2025)
di: Kamo, Naoyuki, et al.
Pubblicazione: (2025)
MOVER: Combining Multiple Meeting Recognition Systems
di: Kamo, Naoyuki, et al.
Pubblicazione: (2025)
di: Kamo, Naoyuki, et al.
Pubblicazione: (2025)
Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder
di: Xie, Yuying, et al.
Pubblicazione: (2024)
di: Xie, Yuying, et al.
Pubblicazione: (2024)
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
di: Peng, Junyi, et al.
Pubblicazione: (2025)
di: Peng, Junyi, et al.
Pubblicazione: (2025)
HiRIS: an Airborne Sonar Sensor with a 1024 Channel Microphone Array for In-Air Acoustic Imaging
di: Laurijssen, Dennis, et al.
Pubblicazione: (2024)
di: Laurijssen, Dennis, et al.
Pubblicazione: (2024)
Error Analysis in a Modular Meeting Transcription System
di: Vieting, Peter, et al.
Pubblicazione: (2025)
di: Vieting, Peter, et al.
Pubblicazione: (2025)
Lessons Learned from the URGENT 2024 Speech Enhancement Challenge
di: Zhang, Wangyou, et al.
Pubblicazione: (2025)
di: Zhang, Wangyou, et al.
Pubblicazione: (2025)
Generative Deep Learning and Signal Processing for Data Augmentation of Cardiac Auscultation Signals: Improving Model Robustness Using Synthetic Audio
di: Abbott, Leigh, et al.
Pubblicazione: (2024)
di: Abbott, Leigh, et al.
Pubblicazione: (2024)
Speech Quality Embeddings for Improved Detection and Classification of Degradations in Speech Signals
di: Kuhlmann, Michael, et al.
Pubblicazione: (2026)
di: Kuhlmann, Michael, et al.
Pubblicazione: (2026)
Investigation of Speaker Representation for Target-Speaker Speech Processing
di: Ashihara, Takanori, et al.
Pubblicazione: (2024)
di: Ashihara, Takanori, et al.
Pubblicazione: (2024)
Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario
di: Wen, Wen, et al.
Pubblicazione: (2024)
di: Wen, Wen, et al.
Pubblicazione: (2024)
Steered Response Power-Based Direction-of-Arrival Estimation Exploiting an Auxiliary Microphone
di: Brümann, Klaus, et al.
Pubblicazione: (2024)
di: Brümann, Klaus, et al.
Pubblicazione: (2024)
BRUDEX Database: Binaural Room Impulse Responses with Uniformly Distributed External Microphones
di: Fejgin, Daniel, et al.
Pubblicazione: (2023)
di: Fejgin, Daniel, et al.
Pubblicazione: (2023)
Toward Universal Speech Enhancement for Diverse Input Conditions
di: Zhang, Wangyou, et al.
Pubblicazione: (2023)
di: Zhang, Wangyou, et al.
Pubblicazione: (2023)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
di: Wang, Kuan-Chen, et al.
Pubblicazione: (2024)
di: Wang, Kuan-Chen, et al.
Pubblicazione: (2024)
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
di: Serre, Thomas, et al.
Pubblicazione: (2026)
di: Serre, Thomas, et al.
Pubblicazione: (2026)
Tool Wear Prediction in CNC Turning Operations using Ultrasonic Microphone Arrays and CNNs
di: Steckel, Jan, et al.
Pubblicazione: (2024)
di: Steckel, Jan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization
di: von Neumann, Thilo, et al.
Pubblicazione: (2023) -
Generic Speech Enhancement with Self-Supervised Representation Space Loss
di: Sato, Hiroshi, et al.
Pubblicazione: (2025) -
Array Geometry-Robust Attention-Based Neural Beamformer for Moving Speakers
di: Tammen, Marvin, et al.
Pubblicazione: (2024) -
Word Error Rate Definitions and Algorithms for Long-Form Multi-talker Speech Recognition
di: von Neumann, Thilo, et al.
Pubblicazione: (2025) -
Loose coupling of spectral and spatial models for multi-channel diarization and enhancement of meetings in dynamic environments
di: Meise, Adrian, et al.
Pubblicazione: (2026)