LocaGen: Sub-Sample Time-Delay Learning for Beam Localization
Fuente:
arXiv
Saved in:
| Main Authors: | Kunwar, Ishaan, Cantor, Henry, Rizzo, Tyler, Qayyum, Ayaan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FullSubNet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement
by: Hao, Xiang, et al.
Published: (2020)
by: Hao, Xiang, et al.
Published: (2020)
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
by: Yan, Haoyin, et al.
Published: (2024)
by: Yan, Haoyin, et al.
Published: (2024)
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
by: Sato, Hiroshi, et al.
Published: (2024)
by: Sato, Hiroshi, et al.
Published: (2024)
Blind Source Separation of Radar Signals in Time Domain Using Deep Learning
by: Hinderer, Sven
Published: (2025)
by: Hinderer, Sven
Published: (2025)
Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
by: Yu, Chin-Yun, et al.
Published: (2022)
by: Yu, Chin-Yun, et al.
Published: (2022)
Interpolation Filter Design for Sample Rate Independent Audio Effect RNNs
by: Carson, Alistair, et al.
Published: (2024)
by: Carson, Alistair, et al.
Published: (2024)
Sample Rate Independent Recurrent Neural Networks for Audio Effects Processing
by: Carson, Alistair, et al.
Published: (2024)
by: Carson, Alistair, et al.
Published: (2024)
Reverberation-based Features for Sound Event Localization and Detection with Distance Estimation
by: Berghi, Davide, et al.
Published: (2025)
by: Berghi, Davide, et al.
Published: (2025)
An Investigation of Time-Frequency Representation Discriminators for High-Fidelity Vocoder
by: Gu, Yicheng, et al.
Published: (2024)
by: Gu, Yicheng, et al.
Published: (2024)
Time-domain sound field estimation using kernel ridge regression
by: Brunnström, Jesper, et al.
Published: (2025)
by: Brunnström, Jesper, et al.
Published: (2025)
RIFT: Entropy-Optimised Fractional Wavelet Constellations for Ideal Time-Frequency Estimation
by: Cozens, James M., et al.
Published: (2025)
by: Cozens, James M., et al.
Published: (2025)
Note-Level Singing Melody Transcription for Time-Aligned Musical Score Generation
by: Kim, Leekyung, et al.
Published: (2025)
by: Kim, Leekyung, et al.
Published: (2025)
Wavelet-Based Time-Frequency Fingerprinting for Feature Extraction of Traditional Irish Music
by: Shore, Noah
Published: (2025)
by: Shore, Noah
Published: (2025)
DDD: A Perceptually Superior Low-Response-Time DNN-based Declipper
by: Yi, Jayeon, et al.
Published: (2024)
by: Yi, Jayeon, et al.
Published: (2024)
Time-of-arrival Estimation and Phase Unwrapping of Head-related Transfer Functions With Integer Linear Programming
by: Yu, Chin-Yun, et al.
Published: (2024)
by: Yu, Chin-Yun, et al.
Published: (2024)
Learning Perceptually Relevant Temporal Envelope Morphing
by: Dixit, Satvik, et al.
Published: (2025)
by: Dixit, Satvik, et al.
Published: (2025)
A Multimodal Data Fusion Attention-Empowered Generative Adversarial Network for Real Time 3D Underwater Sound Speed Field Construction
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
by: Haeb-Umbach, Reinhold, et al.
Published: (2025)
by: Haeb-Umbach, Reinhold, et al.
Published: (2025)
Lessons Learned from the URGENT 2024 Speech Enhancement Challenge
by: Zhang, Wangyou, et al.
Published: (2025)
by: Zhang, Wangyou, et al.
Published: (2025)
Machine Learning in Acoustics: A Review and Open-Source Repository
by: McCarthy, Ryan A., et al.
Published: (2025)
by: McCarthy, Ryan A., et al.
Published: (2025)
AADNet: An End-to-End Deep Learning Model for Auditory Attention Decoding
by: Nguyen, Nhan Duc Thanh, et al.
Published: (2024)
by: Nguyen, Nhan Duc Thanh, et al.
Published: (2024)
Optimal Scalogram for Computational Complexity Reduction in Acoustic Recognition Using Deep Learning
by: Phan, Dang Thoai, et al.
Published: (2025)
by: Phan, Dang Thoai, et al.
Published: (2025)
Bird Vocalization Embedding Extraction Using Self-Supervised Disentangled Representation Learning
by: Shi, Runwu, et al.
Published: (2024)
by: Shi, Runwu, et al.
Published: (2024)
Enhancing Anti-spoofing Countermeasures Robustness through Joint Optimization and Transfer Learning
by: Wang, Yikang, et al.
Published: (2024)
by: Wang, Yikang, et al.
Published: (2024)
Detecting Post-Stroke Aphasia Via Brain Responses to Speech in a Deep Learning Framework
by: De Clercq, Pieter, et al.
Published: (2024)
by: De Clercq, Pieter, et al.
Published: (2024)
Ultrasensitive Textile Strain Sensors Redefine Wearable Silent Speech Interfaces with High Machine Learning Efficiency
by: Tang, Chenyu, et al.
Published: (2023)
by: Tang, Chenyu, et al.
Published: (2023)
Frequency-Based Alignment of EEG and Audio Signals Using Contrastive Learning and SincNet for Auditory Attention Detection
by: Liao, Yuan, et al.
Published: (2025)
by: Liao, Yuan, et al.
Published: (2025)
Generative Deep Learning and Signal Processing for Data Augmentation of Cardiac Auscultation Signals: Improving Model Robustness Using Synthetic Audio
by: Abbott, Leigh, et al.
Published: (2024)
by: Abbott, Leigh, et al.
Published: (2024)
Multi-Source Localization and Data Association for Time-Difference of Arrival Measurements
by: Flood, Gabrielle, et al.
Published: (2024)
by: Flood, Gabrielle, et al.
Published: (2024)
Array-Aware Ambisonics and HRTF Encoding for Binaural Reproduction With Wearable Arrays
by: Gayer, Yhonatan, et al.
Published: (2025)
by: Gayer, Yhonatan, et al.
Published: (2025)
AutoMashup: Automatic Music Mashups Creation
by: Delabaere, Marine, et al.
Published: (2025)
by: Delabaere, Marine, et al.
Published: (2025)
Beamforming in the Reproducing Kernel Domain Based on Spatial Differentiation
by: Iwami, Takahiro, et al.
Published: (2025)
by: Iwami, Takahiro, et al.
Published: (2025)
Tracking of Intermittent and Moving Speakers : Dataset and Metrics
by: Iatariene, Taous, et al.
Published: (2025)
by: Iatariene, Taous, et al.
Published: (2025)
Using Neurogram Similarity Index Measure (NSIM) to Model Hearing Loss and Cochlear Neural Degeneration
by: Cheema, Ahsan J., et al.
Published: (2025)
by: Cheema, Ahsan J., et al.
Published: (2025)
Aliasing-Free Neural Audio Synthesis
by: Gu, Yicheng, et al.
Published: (2025)
by: Gu, Yicheng, et al.
Published: (2025)
Completing Sets of Prototype Transfer Functions for Subspace-based Direction of Arrival Estimation of Multiple Speakers
by: Fejgin, Daniel, et al.
Published: (2025)
by: Fejgin, Daniel, et al.
Published: (2025)
Confidence-Based Self-Training for EMG-to-Speech: Leveraging Synthetic EMG for Robust Modeling
by: Chen, Xiaodan, et al.
Published: (2025)
by: Chen, Xiaodan, et al.
Published: (2025)
DiffAU: Diffusion-Based Ambisonics Upscaling
by: Milstein, Amit, et al.
Published: (2025)
by: Milstein, Amit, et al.
Published: (2025)
Non-locally averaged pruned reassigned spectrograms: a tool for glottal pulse visualization and analysis
by: Griswold, Gabriel J., et al.
Published: (2025)
by: Griswold, Gabriel J., et al.
Published: (2025)
From the perspective of perceptual speech quality: The robustness of frequency bands to noise
by: Fan, Junyi, et al.
Published: (2025)
by: Fan, Junyi, et al.
Published: (2025)
Similar Items
-
FullSubNet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement
by: Hao, Xiang, et al.
Published: (2020) -
LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
by: Yan, Haoyin, et al.
Published: (2024) -
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
by: Sato, Hiroshi, et al.
Published: (2024) -
Blind Source Separation of Radar Signals in Time Domain Using Deep Learning
by: Hinderer, Sven
Published: (2025) -
Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
by: Yu, Chin-Yun, et al.
Published: (2022)