Leveraging Reverberation and Visual Depth Cues for Sound Event Localization and Detection with Distance Estimation
Fuente:
arXiv
Salvato in:
| Autori principali: | Berghi, Davide, Jackson, Philip J. B. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos
di: Berghi, Davide, et al.
Pubblicazione: (2025)
di: Berghi, Davide, et al.
Pubblicazione: (2025)
Spatial and Semantic Embedding Integration for Stereo Sound Event Localization and Detection in Regular Videos
di: Berghi, Davide, et al.
Pubblicazione: (2025)
di: Berghi, Davide, et al.
Pubblicazione: (2025)
Reverberation-based Features for Sound Event Localization and Detection with Distance Estimation
di: Berghi, Davide, et al.
Pubblicazione: (2025)
di: Berghi, Davide, et al.
Pubblicazione: (2025)
ToS: A Team of Specialists ensemble framework for Stereo Sound Event Localization and Detection with distance estimation in Video
di: Berghi, Davide, et al.
Pubblicazione: (2026)
di: Berghi, Davide, et al.
Pubblicazione: (2026)
SoundSil-DS: Deep Denoising and Segmentation of Sound-field Images with Silhouettes
di: Tanigawa, Risako, et al.
Pubblicazione: (2024)
di: Tanigawa, Risako, et al.
Pubblicazione: (2024)
Quality-Controlled Multimodal Emotion Recognition in Conversations with Identity-Based Transfer Learning and MAMBA Fusion
di: Wang, Zanxu, et al.
Pubblicazione: (2025)
di: Wang, Zanxu, et al.
Pubblicazione: (2025)
Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation
di: Neri, Michael, et al.
Pubblicazione: (2026)
di: Neri, Michael, et al.
Pubblicazione: (2026)
Multimodal Biomarkers for Schizophrenia: Towards Individual Symptom Severity Estimation
di: Premananth, Gowtham, et al.
Pubblicazione: (2025)
di: Premananth, Gowtham, et al.
Pubblicazione: (2025)
Multimodal sensor fusion for real-time location-dependent defect detection in laser-directed energy deposition
di: Chen, Lequn, et al.
Pubblicazione: (2023)
di: Chen, Lequn, et al.
Pubblicazione: (2023)
A multi-modal approach for identifying schizophrenia using cross-modal attention
di: Premananth, Gowtham, et al.
Pubblicazione: (2023)
di: Premananth, Gowtham, et al.
Pubblicazione: (2023)
Spoken question answering for visual queries
di: Shabtay, Nimrod, et al.
Pubblicazione: (2025)
di: Shabtay, Nimrod, et al.
Pubblicazione: (2025)
Audio-Visual Approach For Multimodal Concurrent Speaker Detection
di: Eliav, Amit, et al.
Pubblicazione: (2024)
di: Eliav, Amit, et al.
Pubblicazione: (2024)
W4S4: WaLRUS Meets S4 for Long-Range Sequence Modeling
di: Babaei, Hossein, et al.
Pubblicazione: (2025)
di: Babaei, Hossein, et al.
Pubblicazione: (2025)
Fast Swap-Based Element Selection for Multiplication-Free Dimension Reduction
di: Ono, Nobutaka
Pubblicazione: (2026)
di: Ono, Nobutaka
Pubblicazione: (2026)
SaFARi: State-Space Models for Frame-Agnostic Representation
di: Babaei, Hossein, et al.
Pubblicazione: (2025)
di: Babaei, Hossein, et al.
Pubblicazione: (2025)
LiVeAction: a Lightweight, Versatile, and Asymmetric Neural Codec Design for Real-time Operation
di: Jacobellis, Dan, et al.
Pubblicazione: (2026)
di: Jacobellis, Dan, et al.
Pubblicazione: (2026)
End-to-end audio-visual learning for cochlear implant sound coding simulations in noisy environments
di: Lin, Meng-Ping, et al.
Pubblicazione: (2025)
di: Lin, Meng-Ping, et al.
Pubblicazione: (2025)
Leveraging Unlabeled Audio-Visual Data in Speech Emotion Recognition using Knowledge Distillation
di: Pendyala, Varsha, et al.
Pubblicazione: (2025)
di: Pendyala, Varsha, et al.
Pubblicazione: (2025)
U-DREAM: Unsupervised Dereverberation guided by a Reverberation Model
di: Bahrman, Louis, et al.
Pubblicazione: (2025)
di: Bahrman, Louis, et al.
Pubblicazione: (2025)
Exploring Audio-Visual Information Fusion for Sound Event Localization and Detection In Low-Resource Realistic Scenarios
di: Jiang, Ya, et al.
Pubblicazione: (2024)
di: Jiang, Ya, et al.
Pubblicazione: (2024)
Audio-Visual Talker Localization in Video for Spatial Sound Reproduction
di: Berghi, Davide, et al.
Pubblicazione: (2024)
di: Berghi, Davide, et al.
Pubblicazione: (2024)
Livestock feeding behaviour: A review on automated systems for ruminant monitoring
di: Chelotti, José, et al.
Pubblicazione: (2023)
di: Chelotti, José, et al.
Pubblicazione: (2023)
Enhancing Real-World Active Speaker Detection with Multi-Modal Extraction Pre-Training
di: Tao, Ruijie, et al.
Pubblicazione: (2024)
di: Tao, Ruijie, et al.
Pubblicazione: (2024)
Stereo Sound Event Localization and Detection with Onscreen/offscreen Classification
di: Shimada, Kazuki, et al.
Pubblicazione: (2025)
di: Shimada, Kazuki, et al.
Pubblicazione: (2025)
Towards Improving Speaker Distance Estimation through Generative Impulse Response Augmentation
di: Ratnarajah, Anton, et al.
Pubblicazione: (2026)
di: Ratnarajah, Anton, et al.
Pubblicazione: (2026)
WaLRUS: Wavelets for Long-range Representation Using SSMs
di: Babaei, Hossein, et al.
Pubblicazione: (2025)
di: Babaei, Hossein, et al.
Pubblicazione: (2025)
Attentive AV-FusionNet: Audio-Visual Quality Prediction with Hybrid Attention
di: Salaj, Ina, et al.
Pubblicazione: (2025)
di: Salaj, Ina, et al.
Pubblicazione: (2025)
DLIOS: An LLM-Augmented Real-Time Multi-Modal Interactive Enhancement Overlay System for Douyin Live Streaming
di: Wen, Shuide, et al.
Pubblicazione: (2026)
di: Wen, Shuide, et al.
Pubblicazione: (2026)
BUT System Description for CHiME-9 MCoRec Challenge
di: Klement, Dominik, et al.
Pubblicazione: (2026)
di: Klement, Dominik, et al.
Pubblicazione: (2026)
Listening for "You": Enhancing Speech Image Retrieval via Target Speaker Extraction
di: Yang, Wenhao, et al.
Pubblicazione: (2025)
di: Yang, Wenhao, et al.
Pubblicazione: (2025)
Bounds on Agreement between Subjective and Objective Measurements
di: Pieper, Jaden, et al.
Pubblicazione: (2026)
di: Pieper, Jaden, et al.
Pubblicazione: (2026)
Improvement Of Audiovisual Quality Estimation Using A Nonlinear Autoregressive Exogenous Neural Network And Bitstream Parameters
di: Kossi, Koffi, et al.
Pubblicazione: (2024)
di: Kossi, Koffi, et al.
Pubblicazione: (2024)
TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation
di: Kim, Ji-Hoon, et al.
Pubblicazione: (2025)
di: Kim, Ji-Hoon, et al.
Pubblicazione: (2025)
Efficient Face Detection with Audio-Based Region Proposals for Human-Robot Interactions
di: Aris, William, et al.
Pubblicazione: (2023)
di: Aris, William, et al.
Pubblicazione: (2023)
MASSLOC: A Massive Sound Source Localization System based on Direction-of-Arrival Estimation
di: Fischer, Georg K. J., et al.
Pubblicazione: (2025)
di: Fischer, Georg K. J., et al.
Pubblicazione: (2025)
Text-Queried Target Sound Event Localization
di: Zhao, Jinzheng, et al.
Pubblicazione: (2024)
di: Zhao, Jinzheng, et al.
Pubblicazione: (2024)
Reacting like Humans: Incorporating Intrinsic Human Behaviors into NAO through Sound-Based Reactions to Fearful and Shocking Events for Enhanced Sociability
di: Ghadami, Ali, et al.
Pubblicazione: (2023)
di: Ghadami, Ali, et al.
Pubblicazione: (2023)
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization
di: Malard, Hugo, et al.
Pubblicazione: (2024)
di: Malard, Hugo, et al.
Pubblicazione: (2024)
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet
di: Zhong, Zhi, et al.
Pubblicazione: (2025)
di: Zhong, Zhi, et al.
Pubblicazione: (2025)
PersonaCite: VoC-Grounded Interviewable Agentic Synthetic AI Personas for Verifiable User and Design Research
di: Truss, Mario
Pubblicazione: (2026)
di: Truss, Mario
Pubblicazione: (2026)
Documenti analoghi
-
Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos
di: Berghi, Davide, et al.
Pubblicazione: (2025) -
Spatial and Semantic Embedding Integration for Stereo Sound Event Localization and Detection in Regular Videos
di: Berghi, Davide, et al.
Pubblicazione: (2025) -
Reverberation-based Features for Sound Event Localization and Detection with Distance Estimation
di: Berghi, Davide, et al.
Pubblicazione: (2025) -
ToS: A Team of Specialists ensemble framework for Stereo Sound Event Localization and Detection with distance estimation in Video
di: Berghi, Davide, et al.
Pubblicazione: (2026) -
SoundSil-DS: Deep Denoising and Segmentation of Sound-field Images with Silhouettes
di: Tanigawa, Risako, et al.
Pubblicazione: (2024)