Deep Learning Based Stage-wise Two-dimensional Speaker Localization with Large Ad-hoc Microphone Arrays
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Shupei, Feng, Linfeng, Gong, Yijun, Liang, Chengdong, Zhang, Chen, Zhang, Xiao-Lei, Li, Xuelong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LABNet: A Lightweight Attentive Beamforming Network for Ad-hoc Multichannel Microphone Invariant Real-Time Speech Enhancement
von: Yan, Haoyin, et al.
Veröffentlicht: (2025)
von: Yan, Haoyin, et al.
Veröffentlicht: (2025)
Diffusion-Based Adversarial Purification for Speaker Verification
von: Bai, Yibo, et al.
Veröffentlicht: (2023)
von: Bai, Yibo, et al.
Veröffentlicht: (2023)
AudioSpa: Spatializing Sound Events with Text
von: Feng, Linfeng, et al.
Veröffentlicht: (2025)
von: Feng, Linfeng, et al.
Veröffentlicht: (2025)
Blind Localization of Early Room Reflections with Arbitrary Microphone Array
von: Hadadi, Yogev, et al.
Veröffentlicht: (2024)
von: Hadadi, Yogev, et al.
Veröffentlicht: (2024)
Asynchronous Microphone Array Calibration using Hybrid TDOA Information
von: Zhang, Chengjie, et al.
Veröffentlicht: (2024)
von: Zhang, Chengjie, et al.
Veröffentlicht: (2024)
SonicBoom: Contact Localization Using Array of Microphones
von: Lee, Moonyoung, et al.
Veröffentlicht: (2024)
von: Lee, Moonyoung, et al.
Veröffentlicht: (2024)
Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR
von: Ma, Hao, et al.
Veröffentlicht: (2025)
von: Ma, Hao, et al.
Veröffentlicht: (2025)
Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario
von: Wen, Wen, et al.
Veröffentlicht: (2024)
von: Wen, Wen, et al.
Veröffentlicht: (2024)
Single-Microphone Speaker Separation and Voice Activity Detection in Noisy and Reverberant Environments
von: Opochinsky, Renana, et al.
Veröffentlicht: (2024)
von: Opochinsky, Renana, et al.
Veröffentlicht: (2024)
RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization
von: Yang, Bing, et al.
Veröffentlicht: (2024)
von: Yang, Bing, et al.
Veröffentlicht: (2024)
Applying Automatic Differentiation to Optimize Differential Microphone Array Designs
von: Galougah, Siminfar Samakoush, et al.
Veröffentlicht: (2024)
von: Galougah, Siminfar Samakoush, et al.
Veröffentlicht: (2024)
Speaker Contrastive Learning for Source Speaker Tracing
von: Wang, Qing, et al.
Veröffentlicht: (2024)
von: Wang, Qing, et al.
Veröffentlicht: (2024)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
von: Haeb-Umbach, Reinhold, et al.
Veröffentlicht: (2025)
DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided Diffusion Model
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
Your Microphone Array Retains Your Identity: A Robust Voice Liveness Detection System for Smart Speakers
von: Meng, Yan, et al.
Veröffentlicht: (2025)
von: Meng, Yan, et al.
Veröffentlicht: (2025)
Neural Directional Filtering: Far-Field Directivity Control With a Small Microphone Array
von: Wechsler, Julian, et al.
Veröffentlicht: (2024)
von: Wechsler, Julian, et al.
Veröffentlicht: (2024)
VM-UNSSOR: Unsupervised Neural Speech Separation Enhanced by Higher-SNR Virtual Microphone Arrays
von: He, Shulin, et al.
Veröffentlicht: (2025)
von: He, Shulin, et al.
Veröffentlicht: (2025)
High-Fidelity Generative Audio Compression at 0.275kbps
von: Ma, Hao, et al.
Veröffentlicht: (2026)
von: Ma, Hao, et al.
Veröffentlicht: (2026)
Theoretical Framework for the Optimization of Microphone Array Configuration for Humanoid Robot Audition
von: Tourbabin, Vladimir, et al.
Veröffentlicht: (2024)
von: Tourbabin, Vladimir, et al.
Veröffentlicht: (2024)
Feasibility of iMagLS-BSM -- ILD Informed Binaural Signal Matching with Arbitrary Microphone Arrays
von: Berebi, Or, et al.
Veröffentlicht: (2024)
von: Berebi, Or, et al.
Veröffentlicht: (2024)
Performance and Robustness of Signal-Dependent vs. Signal-Independent Binaural Signal Matching with Wearable Microphone Arrays
von: Berger, Ami, et al.
Veröffentlicht: (2024)
von: Berger, Ami, et al.
Veröffentlicht: (2024)
Emotional Styles Hide in Deep Speaker Embeddings: Disentangle Deep Speaker Embeddings for Speaker Clustering
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
von: Lin, Chaohao, et al.
Veröffentlicht: (2025)
Assisted RTF-Vector-Based Binaural Direction of Arrival Estimation Exploiting a Calibrated External Microphone Array
von: Fejgin, Daniel, et al.
Veröffentlicht: (2022)
von: Fejgin, Daniel, et al.
Veröffentlicht: (2022)
BSM-iMagLS: ILD Informed Binaural Signal Matching for Reproduction with Head-Mounted Microphone Arrays
von: Berebi, Or, et al.
Veröffentlicht: (2025)
von: Berebi, Or, et al.
Veröffentlicht: (2025)
Direction of Arrival Estimation Using Microphone Array Processing for Moving Humanoid Robots
von: Tourbabin, Vladimir, et al.
Veröffentlicht: (2024)
von: Tourbabin, Vladimir, et al.
Veröffentlicht: (2024)
Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models
von: Jing, Ruihao, et al.
Veröffentlicht: (2025)
von: Jing, Ruihao, et al.
Veröffentlicht: (2025)
Two-stage Audio-Visual Target Speaker Extraction System for Real-Time Processing On Edge Device
von: Li, Zixuan, et al.
Veröffentlicht: (2025)
von: Li, Zixuan, et al.
Veröffentlicht: (2025)
SpatialEmb: Extract and Encode Spatial Information for 1-Stage Multi-channel Multi-speaker ASR on Arbitrary Microphone Arrays
von: Shao, Yiwen, et al.
Veröffentlicht: (2026)
von: Shao, Yiwen, et al.
Veröffentlicht: (2026)
Spatial Analysis and Synthesis Methods: Subjective and Objective Evaluations Using Various Microphone Arrays in the Auralization of a Critical Listening Room
von: Pawlak, Alan, et al.
Veröffentlicht: (2024)
von: Pawlak, Alan, et al.
Veröffentlicht: (2024)
Array Geometry-Robust Attention-Based Neural Beamformer for Moving Speakers
von: Tammen, Marvin, et al.
Veröffentlicht: (2024)
von: Tammen, Marvin, et al.
Veröffentlicht: (2024)
3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource Applications
von: He, Shulin, et al.
Veröffentlicht: (2023)
von: He, Shulin, et al.
Veröffentlicht: (2023)
Direction Estimation of Sound Sources Using Microphone Arrays and Signal Strength
von: Pour, Mahdi Ali, et al.
Veröffentlicht: (2025)
von: Pour, Mahdi Ali, et al.
Veröffentlicht: (2025)
Gen-A: Generalizing Ambisonics Neural Encoding to Unseen Microphone Arrays
von: Heikkinen, Mikko, et al.
Veröffentlicht: (2025)
von: Heikkinen, Mikko, et al.
Veröffentlicht: (2025)
Adaptive Data Augmentation with NaturalSpeech3 for Far-field Speaker Verification
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
CaSNet: Compress-and-Send Network Based Multi-Device Speech Enhancement Model for Distributed Microphone Arrays
von: Jiang, Chengqian, et al.
Veröffentlicht: (2026)
von: Jiang, Chengqian, et al.
Veröffentlicht: (2026)
MDD: a Mask Diffusion Detector to Protect Speaker Verification Systems from Adversarial Perturbations
von: Bai, Yibo, et al.
Veröffentlicht: (2025)
von: Bai, Yibo, et al.
Veröffentlicht: (2025)
Exploiting an External Microphone for Binaural RTF-Vector-Based Direction of Arrival Estimation for Multiple Speakers
von: Fejgin, Daniel, et al.
Veröffentlicht: (2023)
von: Fejgin, Daniel, et al.
Veröffentlicht: (2023)
SCDNet: Self-supervised Learning Feature-based Speaker Change Detection
von: Li, Yue, et al.
Veröffentlicht: (2024)
von: Li, Yue, et al.
Veröffentlicht: (2024)
Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation Learning
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
Single-Microphone-Based Sound Source Localization for Mobile Robots in Reverberant Environments
von: Wang, Jiang, et al.
Veröffentlicht: (2025)
von: Wang, Jiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LABNet: A Lightweight Attentive Beamforming Network for Ad-hoc Multichannel Microphone Invariant Real-Time Speech Enhancement
von: Yan, Haoyin, et al.
Veröffentlicht: (2025) -
Diffusion-Based Adversarial Purification for Speaker Verification
von: Bai, Yibo, et al.
Veröffentlicht: (2023) -
AudioSpa: Spatializing Sound Events with Text
von: Feng, Linfeng, et al.
Veröffentlicht: (2025) -
Blind Localization of Early Room Reflections with Arbitrary Microphone Array
von: Hadadi, Yogev, et al.
Veröffentlicht: (2024) -
Asynchronous Microphone Array Calibration using Hybrid TDOA Information
von: Zhang, Chengjie, et al.
Veröffentlicht: (2024)