Latent Acoustic Mapping for Direction of Arrival Estimation: A Self-Supervised Approach
Fuente:
arXiv
Salvato in:
| Autori principali: | Roman, Adrian S., Roman, Iran R., Bello, Juan P. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Robust DOA estimation using deep acoustic imaging
di: Roman, Adrian S., et al.
Pubblicazione: (2024)
di: Roman, Adrian S., et al.
Pubblicazione: (2024)
The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS
di: Carone, Brandon James, et al.
Pubblicazione: (2025)
di: Carone, Brandon James, et al.
Pubblicazione: (2025)
Evaluating Multimodal Large Language Models on Core Music Perception Tasks
di: Carone, Brandon James, et al.
Pubblicazione: (2025)
di: Carone, Brandon James, et al.
Pubblicazione: (2025)
Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection
di: Roman, Adrian S., et al.
Pubblicazione: (2025)
di: Roman, Adrian S., et al.
Pubblicazione: (2025)
Multi-Channel Acoustic Echo Cancellation Based on Direction-of-Arrival Estimation
di: Zhao, Fei, et al.
Pubblicazione: (2025)
di: Zhao, Fei, et al.
Pubblicazione: (2025)
Enhanced Sound Event Localization and Detection in Real 360-degree audio-visual soundscapes
di: Roman, Adrian S., et al.
Pubblicazione: (2024)
di: Roman, Adrian S., et al.
Pubblicazione: (2024)
Spatial Scaper: A Library to Simulate and Augment Soundscapes for Sound Event Localization and Detection in Realistic Rooms
di: Roman, Iran R., et al.
Pubblicazione: (2024)
di: Roman, Iran R., et al.
Pubblicazione: (2024)
Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
di: Aboeitta, Ahmed, et al.
Pubblicazione: (2025)
di: Aboeitta, Ahmed, et al.
Pubblicazione: (2025)
Self-supervised Learning for Acoustic Few-Shot Classification
di: Liang, Jingyong, et al.
Pubblicazione: (2024)
di: Liang, Jingyong, et al.
Pubblicazione: (2024)
Robust Target Speaker Direction of Arrival Estimation
di: Li, Zixuan, et al.
Pubblicazione: (2024)
di: Li, Zixuan, et al.
Pubblicazione: (2024)
SoundCompass: Navigating Target Sound Extraction With Effective Directional Clue Integration In Complex Acoustic Scenes
di: Choi, Dayun, et al.
Pubblicazione: (2025)
di: Choi, Dayun, et al.
Pubblicazione: (2025)
Semi-Supervised Self-Learning Enhanced Music Emotion Recognition
di: Sun, Yifu, et al.
Pubblicazione: (2024)
di: Sun, Yifu, et al.
Pubblicazione: (2024)
Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
di: Zhao, Junqi, et al.
Pubblicazione: (2024)
di: Zhao, Junqi, et al.
Pubblicazione: (2024)
Underwater Acoustic Signal Denoising Algorithms: A Survey of the State-of-the-art
di: Gao, Ruobin, et al.
Pubblicazione: (2024)
di: Gao, Ruobin, et al.
Pubblicazione: (2024)
CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
di: Luong, Justin, et al.
Pubblicazione: (2025)
di: Luong, Justin, et al.
Pubblicazione: (2025)
Data-Efficient ASR Personalization for Non-Normative Speech Using an Uncertainty-Based Phoneme Difficulty Score for Guided Sampling
di: Pokel, Niclas, et al.
Pubblicazione: (2025)
di: Pokel, Niclas, et al.
Pubblicazione: (2025)
Leveraging Real Electric Guitar Tones and Effects to Improve Robustness in Guitar Tablature Transcription Modeling
di: Pedroza, Hegel, et al.
Pubblicazione: (2024)
di: Pedroza, Hegel, et al.
Pubblicazione: (2024)
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
di: Yadav, Sarthak, et al.
Pubblicazione: (2024)
di: Yadav, Sarthak, et al.
Pubblicazione: (2024)
Self-Supervised Models for Phoneme Recognition: Applications in Children's Speech for Reading Learning
di: Medin, Lucas Block, et al.
Pubblicazione: (2025)
di: Medin, Lucas Block, et al.
Pubblicazione: (2025)
Layer-wise Investigation of Large-Scale Self-Supervised Music Representation Models
di: Zhou, Yizhi, et al.
Pubblicazione: (2025)
di: Zhou, Yizhi, et al.
Pubblicazione: (2025)
Improving Speech Inversion Through Self-Supervised Embeddings and Enhanced Tract Variables
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2023)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2023)
When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection
di: Zhang, Xiangyu, et al.
Pubblicazione: (2024)
di: Zhang, Xiangyu, et al.
Pubblicazione: (2024)
Guitar-TECHS: An Electric Guitar Dataset Covering Techniques, Musical Excerpts, Chords and Scales Using a Diverse Array of Hardware
di: Pedroza, Hegel, et al.
Pubblicazione: (2025)
di: Pedroza, Hegel, et al.
Pubblicazione: (2025)
Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models
di: Shah, Neil, et al.
Pubblicazione: (2024)
di: Shah, Neil, et al.
Pubblicazione: (2024)
Prototype based Masked Audio Model for Self-Supervised Learning of Sound Event Detection
di: Cai, Pengfei, et al.
Pubblicazione: (2024)
di: Cai, Pengfei, et al.
Pubblicazione: (2024)
Enhancing Speech Emotion Recognition through Segmental Average Pooling of Self-Supervised Learning Features
di: Hyeon, Jonghwan, et al.
Pubblicazione: (2024)
di: Hyeon, Jonghwan, et al.
Pubblicazione: (2024)
Temporal Variability and Multi-Viewed Self-Supervised Representations to Tackle the ASVspoof5 Deepfake Challenge
di: Xie, Yuankun, et al.
Pubblicazione: (2024)
di: Xie, Yuankun, et al.
Pubblicazione: (2024)
On the Effectiveness of Acoustic BPE in Decoder-Only TTS
di: Li, Bohan, et al.
Pubblicazione: (2024)
di: Li, Bohan, et al.
Pubblicazione: (2024)
Acoustic Imaging for UAV Detection: Dense Beamformed Energy Maps and U-Net SELD
di: Rodriguez, Belman Jahir, et al.
Pubblicazione: (2025)
di: Rodriguez, Belman Jahir, et al.
Pubblicazione: (2025)
SCRAPS: Speech Contrastive Representations of Acoustic and Phonetic Spaces
di: Vallés-Pérez, Ivan, et al.
Pubblicazione: (2023)
di: Vallés-Pérez, Ivan, et al.
Pubblicazione: (2023)
AND: Audio Network Dissection for Interpreting Deep Acoustic Models
di: Wu, Tung-Yu, et al.
Pubblicazione: (2024)
di: Wu, Tung-Yu, et al.
Pubblicazione: (2024)
SPEAR: Receiver-to-Receiver Acoustic Neural Warping Field
di: He, Yuhang, et al.
Pubblicazione: (2024)
di: He, Yuhang, et al.
Pubblicazione: (2024)
DEAF: A Benchmark for Diagnostic Evaluation of Acoustic Faithfulness in Audio Language Models
di: Xiong, Jiaqi, et al.
Pubblicazione: (2026)
di: Xiong, Jiaqi, et al.
Pubblicazione: (2026)
Magnitude-Phase Dual-Path Speech Enhancement Network based on Self-Supervised Embedding and Perceptual Contrast Stretch Boosting
di: Mattursun, Alimjan, et al.
Pubblicazione: (2025)
di: Mattursun, Alimjan, et al.
Pubblicazione: (2025)
Play Me Something Icy: Practical Challenges, Explainability and the Semantic Gap in Generative AI Music
di: Allison, Jesse, et al.
Pubblicazione: (2024)
di: Allison, Jesse, et al.
Pubblicazione: (2024)
Translation-Equivariant Self-Supervised Learning for Pitch Estimation with Optimal Transport
di: Torres, Bernardo, et al.
Pubblicazione: (2025)
di: Torres, Bernardo, et al.
Pubblicazione: (2025)
Pronunciation Deviation Analysis Through Voice Cloning and Acoustic Comparison
di: Valdivia, Andrew, et al.
Pubblicazione: (2025)
di: Valdivia, Andrew, et al.
Pubblicazione: (2025)
AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis
di: Cao, Yubing, et al.
Pubblicazione: (2025)
di: Cao, Yubing, et al.
Pubblicazione: (2025)
AquaSignal: An Integrated Framework for Robust Underwater Acoustic Analysis
di: Panteli, Eirini, et al.
Pubblicazione: (2025)
di: Panteli, Eirini, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Robust DOA estimation using deep acoustic imaging
di: Roman, Adrian S., et al.
Pubblicazione: (2024) -
The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS
di: Carone, Brandon James, et al.
Pubblicazione: (2025) -
Evaluating Multimodal Large Language Models on Core Music Perception Tasks
di: Carone, Brandon James, et al.
Pubblicazione: (2025) -
Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection
di: Roman, Adrian S., et al.
Pubblicazione: (2025) -
Multi-Channel Acoustic Echo Cancellation Based on Direction-of-Arrival Estimation
di: Zhao, Fei, et al.
Pubblicazione: (2025)