FUN-SSL: Full-band Layer Followed by U-Net with Narrow-band Layers for Multiple Moving Sound Source Localization
Fuente:
arXiv
Saved in:
| Main Authors: | Choi, Yuseon, Kim, Hyeonseung, Jun, Jewoo, Shin, Jong Won |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlowSE: Flow Matching-based Speech Enhancement
by: Lee, Seonggyu, et al.
Published: (2025)
by: Lee, Seonggyu, et al.
Published: (2025)
MASSLOC: A Massive Sound Source Localization System based on Direction-of-Arrival Estimation
by: Fischer, Georg K. J., et al.
Published: (2025)
by: Fischer, Georg K. J., et al.
Published: (2025)
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
by: Yan, Haoyin, et al.
Published: (2024)
by: Yan, Haoyin, et al.
Published: (2024)
Speech Enhancement based on cascaded two flows
by: Lee, Seonggyu, et al.
Published: (2025)
by: Lee, Seonggyu, et al.
Published: (2025)
Efficient Personalization of Amplification in Hearing Aids via Multi-band Bayesian Machine Learning
by: Ni, Aoxin, et al.
Published: (2024)
by: Ni, Aoxin, et al.
Published: (2024)
DeepFilterGAN: A Full-band Real-time Speech Enhancement System with GAN-based Stochastic Regeneration
by: Serbest, Sanberk, et al.
Published: (2025)
by: Serbest, Sanberk, et al.
Published: (2025)
From the perspective of perceptual speech quality: The robustness of frequency bands to noise
by: Fan, Junyi, et al.
Published: (2025)
by: Fan, Junyi, et al.
Published: (2025)
Predictive Directional Selective Fixed-Filter Active Noise Control for Moving Sources via a Convolutional Recurrent Neural Network
by: Wang, Boxiang, et al.
Published: (2026)
by: Wang, Boxiang, et al.
Published: (2026)
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
by: Liu, Xiaoyu, et al.
Published: (2024)
by: Liu, Xiaoyu, et al.
Published: (2024)
Sub-band and Full-band Interactive U-Net with DPRNN for Demixing Cross-talk Stereo Music
by: Yin, Han, et al.
Published: (2024)
by: Yin, Han, et al.
Published: (2024)
Exploring Audio-Visual Information Fusion for Sound Event Localization and Detection In Low-Resource Realistic Scenarios
by: Jiang, Ya, et al.
Published: (2024)
by: Jiang, Ya, et al.
Published: (2024)
FullSubNet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement
by: Hao, Xiang, et al.
Published: (2020)
by: Hao, Xiang, et al.
Published: (2020)
Time Difference of Arrival Source Localization: Exact Linear Solutions for the General 3D Problem
by: Inamdar, Niraj K.
Published: (2025)
by: Inamdar, Niraj K.
Published: (2025)
Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?
by: Sinha, Abhijit, et al.
Published: (2025)
by: Sinha, Abhijit, et al.
Published: (2025)
MaskSR: Masked Language Model for Full-band Speech Restoration
by: Li, Xu, et al.
Published: (2024)
by: Li, Xu, et al.
Published: (2024)
A Novel Numerical Method for Relaxing the Minimal Configurations of TOA-Based Joint Sensors and Sources Localization
by: Cao, Faxian, et al.
Published: (2024)
by: Cao, Faxian, et al.
Published: (2024)
String Sound Synthesizer on GPU-accelerated Finite Difference Scheme
by: Lee, Jin Woo, et al.
Published: (2023)
by: Lee, Jin Woo, et al.
Published: (2023)
Reverberation-based Features for Sound Event Localization and Detection with Distance Estimation
by: Berghi, Davide, et al.
Published: (2025)
by: Berghi, Davide, et al.
Published: (2025)
Classification of Adventitious Sounds Combining Cochleogram and Vision Transformers
by: Mang, Loredana Daria, et al.
Published: (2024)
by: Mang, Loredana Daria, et al.
Published: (2024)
Sound field estimation with moving microphones using kernel ridge regression
by: Brunnström, Jesper, et al.
Published: (2025)
by: Brunnström, Jesper, et al.
Published: (2025)
Zero-Shot KWS for Children's Speech using Layer-Wise Features from SSL Models
by: Kutum, Subham, et al.
Published: (2025)
by: Kutum, Subham, et al.
Published: (2025)
Robust Fixed-Filter Sound Zone Control with Audio-Based Position Tracking
by: Bhattacharjee, Sankha Subhra, et al.
Published: (2024)
by: Bhattacharjee, Sankha Subhra, et al.
Published: (2024)
A Zero-Shot Physics-Informed Dictionary Learning Approach for Sound Field Reconstruction
by: Damiano, Stefano, et al.
Published: (2024)
by: Damiano, Stefano, et al.
Published: (2024)
Tracking of Intermittent and Moving Speakers : Dataset and Metrics
by: Iatariene, Taous, et al.
Published: (2025)
by: Iatariene, Taous, et al.
Published: (2025)
SyncNet: correlating objective for time delay estimation in audio signals
by: Raina, Akshay, et al.
Published: (2022)
by: Raina, Akshay, et al.
Published: (2022)
SUNAC: Source-aware Unified Neural Audio Codec
by: Aihara, Ryo, et al.
Published: (2025)
by: Aihara, Ryo, et al.
Published: (2025)
ILD-VIT: A Unified Vision Transformer Architecture for Detection of Interstitial Lung Disease from Respiratory Sounds
by: Hota, Soubhagya Ranjan, et al.
Published: (2025)
by: Hota, Soubhagya Ranjan, et al.
Published: (2025)
DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing
by: Lee, Dongheon, et al.
Published: (2023)
by: Lee, Dongheon, et al.
Published: (2023)
Multi-Source Position and Direction-of-Arrival Estimation Based on Euclidean Distance Matrices
by: Brümann, Klaus, et al.
Published: (2025)
by: Brümann, Klaus, et al.
Published: (2025)
Binaural Localization Model for Speech in Noise
by: Tokala, Vikas, et al.
Published: (2025)
by: Tokala, Vikas, et al.
Published: (2025)
ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
by: Chen, Yafeng, et al.
Published: (2024)
by: Chen, Yafeng, et al.
Published: (2024)
3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization
by: Chen, Yafeng, et al.
Published: (2024)
by: Chen, Yafeng, et al.
Published: (2024)
Speakers Localization Using Batch EM In Unfolding Neural Network
by: Veler, Rina, et al.
Published: (2026)
by: Veler, Rina, et al.
Published: (2026)
QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model
by: Chen, Shaowen, et al.
Published: (2025)
by: Chen, Shaowen, et al.
Published: (2025)
FlowDec: A flow-based full-band general audio codec with high perceptual quality
by: Welker, Simon, et al.
Published: (2025)
by: Welker, Simon, et al.
Published: (2025)
Musical Score Following using Statistical Inference
by: Cowley, Josephine
Published: (2025)
by: Cowley, Josephine
Published: (2025)
Tracking of Spatially Dynamic Room Impulse Responses Along Locally Linearized Trajectories
by: MacWilliam, Kathleen, et al.
Published: (2025)
by: MacWilliam, Kathleen, et al.
Published: (2025)
Blind Estimation of Sub-band Acoustic Parameters from Ambisonics Recordings using Spectro-Spatial Covariance Features
by: Meng, Hanyu, et al.
Published: (2024)
by: Meng, Hanyu, et al.
Published: (2024)
Permutation Invariant Recurrent Neural Networks for Sound Source Tracking Applications
by: Diaz-Guerra, David, et al.
Published: (2023)
by: Diaz-Guerra, David, et al.
Published: (2023)
Dynamic Prediction of Full-Ocean Depth SSP by Hierarchical LSTM: An Experimental Result
by: Lu, Jiajun, et al.
Published: (2023)
by: Lu, Jiajun, et al.
Published: (2023)
Similar Items
-
FlowSE: Flow Matching-based Speech Enhancement
by: Lee, Seonggyu, et al.
Published: (2025) -
MASSLOC: A Massive Sound Source Localization System based on Direction-of-Arrival Estimation
by: Fischer, Georg K. J., et al.
Published: (2025) -
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
by: Yan, Haoyin, et al.
Published: (2024) -
Speech Enhancement based on cascaded two flows
by: Lee, Seonggyu, et al.
Published: (2025) -
Efficient Personalization of Amplification in Hearing Aids via Multi-band Bayesian Machine Learning
by: Ni, Aoxin, et al.
Published: (2024)