RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Bing, Quan, Changsheng, Wang, Yabo, Wang, Pengyu, Yang, Yujie, Fang, Ying, Shao, Nian, Bu, Hui, Xu, Xin, Li, Xiaofei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multichannel Long-Term Streaming Neural Speech Enhancement for Static and Moving Speakers
von: Quan, Changsheng, et al.
Veröffentlicht: (2024)
von: Quan, Changsheng, et al.
Veröffentlicht: (2024)
CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR
von: Shao, Nian, et al.
Veröffentlicht: (2025)
von: Shao, Nian, et al.
Veröffentlicht: (2025)
IPDnet: A Universal Direct-Path IPD Estimation Network for Sound Source Localization
von: Wang, Yabo, et al.
Veröffentlicht: (2024)
von: Wang, Yabo, et al.
Veröffentlicht: (2024)
Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement
von: Yang, Yujie, et al.
Veröffentlicht: (2025)
von: Yang, Yujie, et al.
Veröffentlicht: (2025)
Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario
von: Wen, Wen, et al.
Veröffentlicht: (2024)
von: Wen, Wen, et al.
Veröffentlicht: (2024)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
LABNet: A Lightweight Attentive Beamforming Network for Ad-hoc Multichannel Microphone Invariant Real-Time Speech Enhancement
von: Yan, Haoyin, et al.
Veröffentlicht: (2025)
von: Yan, Haoyin, et al.
Veröffentlicht: (2025)
VM-UNSSOR: Unsupervised Neural Speech Separation Enhanced by Higher-SNR Virtual Microphone Arrays
von: He, Shulin, et al.
Veröffentlicht: (2025)
von: He, Shulin, et al.
Veröffentlicht: (2025)
Asynchronous Microphone Array Calibration using Hybrid TDOA Information
von: Zhang, Chengjie, et al.
Veröffentlicht: (2024)
von: Zhang, Chengjie, et al.
Veröffentlicht: (2024)
Fine-tune the pretrained ATST model for sound event detection
von: Shao, Nian, et al.
Veröffentlicht: (2023)
von: Shao, Nian, et al.
Veröffentlicht: (2023)
CaSNet: Compress-and-Send Network Based Multi-Device Speech Enhancement Model for Distributed Microphone Arrays
von: Jiang, Chengqian, et al.
Veröffentlicht: (2026)
von: Jiang, Chengqian, et al.
Veröffentlicht: (2026)
Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module
von: Cui, Zhongjian, et al.
Veröffentlicht: (2025)
von: Cui, Zhongjian, et al.
Veröffentlicht: (2025)
Applying Automatic Differentiation to Optimize Differential Microphone Array Designs
von: Galougah, Siminfar Samakoush, et al.
Veröffentlicht: (2024)
von: Galougah, Siminfar Samakoush, et al.
Veröffentlicht: (2024)
Blind Localization of Early Room Reflections with Arbitrary Microphone Array
von: Hadadi, Yogev, et al.
Veröffentlicht: (2024)
von: Hadadi, Yogev, et al.
Veröffentlicht: (2024)
Self-Supervised Learning of Spatial Acoustic Representation with Cross-Channel Signal Reconstruction and Multi-Channel Conformer
von: Yang, Bing, et al.
Veröffentlicht: (2023)
von: Yang, Bing, et al.
Veröffentlicht: (2023)
SonicBoom: Contact Localization Using Array of Microphones
von: Lee, Moonyoung, et al.
Veröffentlicht: (2024)
von: Lee, Moonyoung, et al.
Veröffentlicht: (2024)
A Two-Stage Framework in Cross-Spectrum Domain for Real-Time Speech Enhancement
von: Zhang, Yuewei, et al.
Veröffentlicht: (2024)
von: Zhang, Yuewei, et al.
Veröffentlicht: (2024)
Head Orientation Estimation with Distributed Microphones Using Speech Radiation Patterns
von: Müller, Kaspar, et al.
Veröffentlicht: (2023)
von: Müller, Kaspar, et al.
Veröffentlicht: (2023)
FullSubNet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement
von: Hao, Xiang, et al.
Veröffentlicht: (2020)
von: Hao, Xiang, et al.
Veröffentlicht: (2020)
Neural Directional Filtering: Far-Field Directivity Control With a Small Microphone Array
von: Wechsler, Julian, et al.
Veröffentlicht: (2024)
von: Wechsler, Julian, et al.
Veröffentlicht: (2024)
Modeling of Speech-dependent Own Voice Transfer Characteristics for Hearables with In-ear Microphones
von: Ohlenbusch, Mattes, et al.
Veröffentlicht: (2023)
von: Ohlenbusch, Mattes, et al.
Veröffentlicht: (2023)
Speech-dependent Modeling of Own Voice Transfer Characteristics for In-ear Microphones in Hearables
von: Ohlenbusch, Mattes, et al.
Veröffentlicht: (2023)
von: Ohlenbusch, Mattes, et al.
Veröffentlicht: (2023)
Leveraging Spatial Cues from Cochlear Implant Microphones to Efficiently Enhance Speech Separation in Real-World Listening Scenes
von: Olalere, Feyisayo, et al.
Veröffentlicht: (2025)
von: Olalere, Feyisayo, et al.
Veröffentlicht: (2025)
SingVERSE: A Diverse, Real-World Benchmark for Singing Voice Enhancement
von: Jiang, Shaohan, et al.
Veröffentlicht: (2025)
von: Jiang, Shaohan, et al.
Veröffentlicht: (2025)
Theoretical Framework for the Optimization of Microphone Array Configuration for Humanoid Robot Audition
von: Tourbabin, Vladimir, et al.
Veröffentlicht: (2024)
von: Tourbabin, Vladimir, et al.
Veröffentlicht: (2024)
WenetSpeech-Wu: Datasets, Benchmarks, and Models for a Unified Chinese Wu Dialect Speech Processing Ecosystem
von: Wang, Chengyou, et al.
Veröffentlicht: (2026)
von: Wang, Chengyou, et al.
Veröffentlicht: (2026)
Feasibility of iMagLS-BSM -- ILD Informed Binaural Signal Matching with Arbitrary Microphone Arrays
von: Berebi, Or, et al.
Veröffentlicht: (2024)
von: Berebi, Or, et al.
Veröffentlicht: (2024)
Performance and Robustness of Signal-Dependent vs. Signal-Independent Binaural Signal Matching with Wearable Microphone Arrays
von: Berger, Ami, et al.
Veröffentlicht: (2024)
von: Berger, Ami, et al.
Veröffentlicht: (2024)
Enroll-on-Wakeup: A First Comparative Study of Target Speech Extraction for Seamless Interaction in Real Noisy Human-Machine Dialogue Scenarios
von: Yang, Yiming, et al.
Veröffentlicht: (2026)
von: Yang, Yiming, et al.
Veröffentlicht: (2026)
WildElder: A Chinese Elderly Speech Dataset from the Wild with Fine-Grained Manual Annotations
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
DroFiT: A Lightweight Band-fused Frequency Attention Toward Real-time UAV Speech Enhancement
von: Lee, Jeongmin, et al.
Veröffentlicht: (2025)
von: Lee, Jeongmin, et al.
Veröffentlicht: (2025)
ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
RCT: Random Consistency Training for Semi-supervised Sound Event Detection
von: Shao, Nian, et al.
Veröffentlicht: (2021)
von: Shao, Nian, et al.
Veröffentlicht: (2021)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
The Universal Personalizer: Few-Shot Dysarthric Speech Recognition via Meta-Learning
von: Agarwal, Dhruuv, et al.
Veröffentlicht: (2025)
von: Agarwal, Dhruuv, et al.
Veröffentlicht: (2025)
A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
von: Wang, Jingyuan, et al.
Veröffentlicht: (2024)
von: Wang, Jingyuan, et al.
Veröffentlicht: (2024)
UniArray: Unified Spectral-Spatial Modeling for Array-Geometry-Agnostic Speech Separation
von: Chen, Weiguang, et al.
Veröffentlicht: (2025)
von: Chen, Weiguang, et al.
Veröffentlicht: (2025)
Confidence-based Filtering for Speech Dataset Curation with Generative Speech Enhancement Using Discrete Tokens
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2026)
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2026)
Direction of Arrival Estimation Using Microphone Array Processing for Moving Humanoid Robots
von: Tourbabin, Vladimir, et al.
Veröffentlicht: (2024)
von: Tourbabin, Vladimir, et al.
Veröffentlicht: (2024)
Deep Learning Based Stage-wise Two-dimensional Speaker Localization with Large Ad-hoc Microphone Arrays
von: Liu, Shupei, et al.
Veröffentlicht: (2022)
von: Liu, Shupei, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Multichannel Long-Term Streaming Neural Speech Enhancement for Static and Moving Speakers
von: Quan, Changsheng, et al.
Veröffentlicht: (2024) -
CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR
von: Shao, Nian, et al.
Veröffentlicht: (2025) -
IPDnet: A Universal Direct-Path IPD Estimation Network for Sound Source Localization
von: Wang, Yabo, et al.
Veröffentlicht: (2024) -
Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement
von: Yang, Yujie, et al.
Veröffentlicht: (2025) -
Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario
von: Wen, Wen, et al.
Veröffentlicht: (2024)