IPDnet2: an efficient and improved inter-channel phase difference estimation network for sound source localization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yabo, Yang, Bing, Li, Xiaofei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
IPDnet: A Universal Direct-Path IPD Estimation Network for Sound Source Localization
von: Wang, Yabo, et al.
Veröffentlicht: (2024)
von: Wang, Yabo, et al.
Veröffentlicht: (2024)
Adaptive high-precision sound source localization at low frequencies based on convolutional neural network
von: Ma, Wenbo, et al.
Veröffentlicht: (2024)
von: Ma, Wenbo, et al.
Veröffentlicht: (2024)
Self-Supervised Learning of Spatial Acoustic Representation with Cross-Channel Signal Reconstruction and Multi-Channel Conformer
von: Yang, Bing, et al.
Veröffentlicht: (2023)
von: Yang, Bing, et al.
Veröffentlicht: (2023)
The Neural-SRP method for positional sound source localization
von: Grinstein, Eric, et al.
Veröffentlicht: (2024)
von: Grinstein, Eric, et al.
Veröffentlicht: (2024)
Fine-tune the pretrained ATST model for sound event detection
von: Shao, Nian, et al.
Veröffentlicht: (2023)
von: Shao, Nian, et al.
Veröffentlicht: (2023)
Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation
von: Vo, Quoc Thinh, et al.
Veröffentlicht: (2025)
von: Vo, Quoc Thinh, et al.
Veröffentlicht: (2025)
RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization
von: Yang, Bing, et al.
Veröffentlicht: (2024)
von: Yang, Bing, et al.
Veröffentlicht: (2024)
Binaural sound source localization using a hybrid time and frequency domain model
von: Geva, Gil, et al.
Veröffentlicht: (2024)
von: Geva, Gil, et al.
Veröffentlicht: (2024)
Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement
von: Yang, Yujie, et al.
Veröffentlicht: (2025)
von: Yang, Yujie, et al.
Veröffentlicht: (2025)
Exterior sound field estimation based on physics-constrained kernel
von: Ribeiro, Juliano G. C., et al.
Veröffentlicht: (2026)
von: Ribeiro, Juliano G. C., et al.
Veröffentlicht: (2026)
Exploiting spatial diversity for increasing the robustness of sound source localization systems against reverberation
von: Garcia-Barrios, Guillermo, et al.
Veröffentlicht: (2024)
von: Garcia-Barrios, Guillermo, et al.
Veröffentlicht: (2024)
In situ sound absorption estimation with the discrete complex image source method
von: Brandao, Eric, et al.
Veröffentlicht: (2024)
von: Brandao, Eric, et al.
Veröffentlicht: (2024)
Human-mimetic binaural ear design and sound source direction estimation for task realization of musculoskeletal humanoids
von: Omura, Yusuke, et al.
Veröffentlicht: (2024)
von: Omura, Yusuke, et al.
Veröffentlicht: (2024)
Representational learning for an anomalous sound detection system with source separation model
von: Shin, Seunghyeon, et al.
Veröffentlicht: (2024)
von: Shin, Seunghyeon, et al.
Veröffentlicht: (2024)
Kernel ridge regression based sound field estimation using a rigid spherical microphone array
von: Matsuda, Ryo, et al.
Veröffentlicht: (2025)
von: Matsuda, Ryo, et al.
Veröffentlicht: (2025)
Rec-RIR: Monaural Blind Room Impulse Response Identification via DNN-based Reverberant Speech Reconstruction in STFT Domain
von: Wang, Pengyu, et al.
Veröffentlicht: (2025)
von: Wang, Pengyu, et al.
Veröffentlicht: (2025)
Stereo sound event localization and detection based on PSELDnet pretraining and BiMamba sequence modeling
von: Gao, Wenmiao, et al.
Veröffentlicht: (2025)
von: Gao, Wenmiao, et al.
Veröffentlicht: (2025)
Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders
von: Dai, Lingling, et al.
Veröffentlicht: (2025)
von: Dai, Lingling, et al.
Veröffentlicht: (2025)
Interaural time difference loss for binaural target sound extraction
von: Hernandez-Olivan, Carlos, et al.
Veröffentlicht: (2024)
von: Hernandez-Olivan, Carlos, et al.
Veröffentlicht: (2024)
Text2Move: Text-to-moving sound generation via trajectory prediction and temporal alignment
von: Liu, Yunyi, et al.
Veröffentlicht: (2025)
von: Liu, Yunyi, et al.
Veröffentlicht: (2025)
Reference Channel Selection by Multi-Channel Masking for End-to-End Multi-Channel Speech Enhancement
von: Dai, Wang, et al.
Veröffentlicht: (2024)
von: Dai, Wang, et al.
Veröffentlicht: (2024)
A Composite Predictive-Generative Approach to Monaural Universal Speech Enhancement
von: Zhang, Jie, et al.
Veröffentlicht: (2025)
von: Zhang, Jie, et al.
Veröffentlicht: (2025)
Multispecies bird sound recognition using a fully convolutional neural network
von: García-Ordás, María Teresa, et al.
Veröffentlicht: (2024)
von: García-Ordás, María Teresa, et al.
Veröffentlicht: (2024)
Mel-FullSubNet: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
Time-domain sound field estimation using kernel ridge regression
von: Brunnström, Jesper, et al.
Veröffentlicht: (2025)
von: Brunnström, Jesper, et al.
Veröffentlicht: (2025)
Controlling the Parameterized Multi-channel Wiener Filter using a tiny neural network
von: Grinstein, Eric, et al.
Veröffentlicht: (2025)
von: Grinstein, Eric, et al.
Veröffentlicht: (2025)
SLM-S2ST: A multimodal language model for direct speech-to-speech translation
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
The effect of self-motion and room familiarity on sound source localization in virtual environments
von: Isserstedt, Niklas, et al.
Veröffentlicht: (2024)
von: Isserstedt, Niklas, et al.
Veröffentlicht: (2024)
LS-EEND: Long-Form Streaming End-to-End Neural Diarization with Online Attractor Extraction
von: Liang, Di, et al.
Veröffentlicht: (2024)
von: Liang, Di, et al.
Veröffentlicht: (2024)
Multichannel Long-Term Streaming Neural Speech Enhancement for Static and Moving Speakers
von: Quan, Changsheng, et al.
Veröffentlicht: (2024)
von: Quan, Changsheng, et al.
Veröffentlicht: (2024)
SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios
von: Li, Kai, et al.
Veröffentlicht: (2024)
von: Li, Kai, et al.
Veröffentlicht: (2024)
Isolation performance metrics for personal sound zone reproduction systems
von: Qiao, Yue, et al.
Veröffentlicht: (2022)
von: Qiao, Yue, et al.
Veröffentlicht: (2022)
Low algorithmic delay implementation of convolutional beamformer for online joint source separation and dereverberation
von: Mo, Kaien, et al.
Veröffentlicht: (2024)
von: Mo, Kaien, et al.
Veröffentlicht: (2024)
Investigation of Speech and Noise Latent Representations in Single-channel VAE-based Speech Enhancement
von: Li, Jiatong, et al.
Veröffentlicht: (2025)
von: Li, Jiatong, et al.
Veröffentlicht: (2025)
The feasibility of sound zone control using an array of parametric array loudspeakers
von: Zhuang, Tao, et al.
Veröffentlicht: (2024)
von: Zhuang, Tao, et al.
Veröffentlicht: (2024)
I-DCCRN-VAE: An Improved Deep Representation Learning Framework for Complex VAE-based Single-channel Speech Enhancement
von: Li, Jiatong, et al.
Veröffentlicht: (2025)
von: Li, Jiatong, et al.
Veröffentlicht: (2025)
Sound event localization and detection based on crnn using rectangular filters and channel rotation data augmentation
von: Ronchini, Francesca, et al.
Veröffentlicht: (2020)
von: Ronchini, Francesca, et al.
Veröffentlicht: (2020)
Acoustic source localization in the spherical harmonics domain exploiting low-rank approximations
von: Cobos, Maximo, et al.
Veröffentlicht: (2023)
von: Cobos, Maximo, et al.
Veröffentlicht: (2023)
Differentiable physics for sound field reconstruction
von: Verburg, Samuel A., et al.
Veröffentlicht: (2025)
von: Verburg, Samuel A., et al.
Veröffentlicht: (2025)
Efficient training strategies for natural sounding speech synthesis and speaker adaptation based on FastPitch
von: Răgman, Teodora, et al.
Veröffentlicht: (2024)
von: Răgman, Teodora, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
IPDnet: A Universal Direct-Path IPD Estimation Network for Sound Source Localization
von: Wang, Yabo, et al.
Veröffentlicht: (2024) -
Adaptive high-precision sound source localization at low frequencies based on convolutional neural network
von: Ma, Wenbo, et al.
Veröffentlicht: (2024) -
Self-Supervised Learning of Spatial Acoustic Representation with Cross-Channel Signal Reconstruction and Multi-Channel Conformer
von: Yang, Bing, et al.
Veröffentlicht: (2023) -
The Neural-SRP method for positional sound source localization
von: Grinstein, Eric, et al.
Veröffentlicht: (2024) -
Fine-tune the pretrained ATST model for sound event detection
von: Shao, Nian, et al.
Veröffentlicht: (2023)