Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Berghi, Davide, Jackson, Philip J. B. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spatial and Semantic Embedding Integration for Stereo Sound Event Localization and Detection in Regular Videos
by: Berghi, Davide, et al.
Published: (2025)
by: Berghi, Davide, et al.
Published: (2025)
ToS: A Team of Specialists ensemble framework for Stereo Sound Event Localization and Detection with distance estimation in Video
by: Berghi, Davide, et al.
Published: (2026)
by: Berghi, Davide, et al.
Published: (2026)
Leveraging Reverberation and Visual Depth Cues for Sound Event Localization and Detection with Distance Estimation
by: Berghi, Davide, et al.
Published: (2024)
by: Berghi, Davide, et al.
Published: (2024)
Reverberation-based Features for Sound Event Localization and Detection with Distance Estimation
by: Berghi, Davide, et al.
Published: (2025)
by: Berghi, Davide, et al.
Published: (2025)
Quality-Controlled Multimodal Emotion Recognition in Conversations with Identity-Based Transfer Learning and MAMBA Fusion
by: Wang, Zanxu, et al.
Published: (2025)
by: Wang, Zanxu, et al.
Published: (2025)
Multimodal Biomarkers for Schizophrenia: Towards Individual Symptom Severity Estimation
by: Premananth, Gowtham, et al.
Published: (2025)
by: Premananth, Gowtham, et al.
Published: (2025)
W4S4: WaLRUS Meets S4 for Long-Range Sequence Modeling
by: Babaei, Hossein, et al.
Published: (2025)
by: Babaei, Hossein, et al.
Published: (2025)
Fast Swap-Based Element Selection for Multiplication-Free Dimension Reduction
by: Ono, Nobutaka
Published: (2026)
by: Ono, Nobutaka
Published: (2026)
SaFARi: State-Space Models for Frame-Agnostic Representation
by: Babaei, Hossein, et al.
Published: (2025)
by: Babaei, Hossein, et al.
Published: (2025)
LiVeAction: a Lightweight, Versatile, and Asymmetric Neural Codec Design for Real-time Operation
by: Jacobellis, Dan, et al.
Published: (2026)
by: Jacobellis, Dan, et al.
Published: (2026)
Livestock feeding behaviour: A review on automated systems for ruminant monitoring
by: Chelotti, José, et al.
Published: (2023)
by: Chelotti, José, et al.
Published: (2023)
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization
by: Malard, Hugo, et al.
Published: (2024)
by: Malard, Hugo, et al.
Published: (2024)
SoundSil-DS: Deep Denoising and Segmentation of Sound-field Images with Silhouettes
by: Tanigawa, Risako, et al.
Published: (2024)
by: Tanigawa, Risako, et al.
Published: (2024)
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet
by: Zhong, Zhi, et al.
Published: (2025)
by: Zhong, Zhi, et al.
Published: (2025)
Multimodal sensor fusion for real-time location-dependent defect detection in laser-directed energy deposition
by: Chen, Lequn, et al.
Published: (2023)
by: Chen, Lequn, et al.
Published: (2023)
WaLRUS: Wavelets for Long-range Representation Using SSMs
by: Babaei, Hossein, et al.
Published: (2025)
by: Babaei, Hossein, et al.
Published: (2025)
Video Soundtrack Generation by Aligning Emotions and Temporal Boundaries
by: Sulun, Serkan, et al.
Published: (2025)
by: Sulun, Serkan, et al.
Published: (2025)
Reacting like Humans: Incorporating Intrinsic Human Behaviors into NAO through Sound-Based Reactions to Fearful and Shocking Events for Enhanced Sociability
by: Ghadami, Ali, et al.
Published: (2023)
by: Ghadami, Ali, et al.
Published: (2023)
A multi-modal approach for identifying schizophrenia using cross-modal attention
by: Premananth, Gowtham, et al.
Published: (2023)
by: Premananth, Gowtham, et al.
Published: (2023)
Audio-Visual Talker Localization in Video for Spatial Sound Reproduction
by: Berghi, Davide, et al.
Published: (2024)
by: Berghi, Davide, et al.
Published: (2024)
Learned Compression for Compressed Learning
by: Jacobellis, Dan, et al.
Published: (2024)
by: Jacobellis, Dan, et al.
Published: (2024)
SIM: Surface-based fMRI Analysis for Inter-Subject Multimodal Decoding from Movie-Watching Experiments
by: Dahan, Simon, et al.
Published: (2025)
by: Dahan, Simon, et al.
Published: (2025)
T-FOLEY: A Controllable Waveform-Domain Diffusion Model for Temporal-Event-Guided Foley Sound Synthesis
by: Chung, Yoonjin, et al.
Published: (2024)
by: Chung, Yoonjin, et al.
Published: (2024)
Spoken question answering for visual queries
by: Shabtay, Nimrod, et al.
Published: (2025)
by: Shabtay, Nimrod, et al.
Published: (2025)
Manikin-Recorded Cardiopulmonary Sounds Dataset Using Digital Stethoscope
by: Torabi, Yasaman, et al.
Published: (2024)
by: Torabi, Yasaman, et al.
Published: (2024)
MEMS and ECM Sensor Technologies for Cardiorespiratory Sound Monitoring - A Comprehensive Review
by: Torabi, Yasaman, et al.
Published: (2024)
by: Torabi, Yasaman, et al.
Published: (2024)
Stereo Sound Event Localization and Detection with Onscreen/offscreen Classification
by: Shimada, Kazuki, et al.
Published: (2025)
by: Shimada, Kazuki, et al.
Published: (2025)
Leveraging Unlabeled Audio-Visual Data in Speech Emotion Recognition using Knowledge Distillation
by: Pendyala, Varsha, et al.
Published: (2025)
by: Pendyala, Varsha, et al.
Published: (2025)
RestoreGrad: Signal Restoration Using Conditional Denoising Diffusion Models with Jointly Learned Prior
by: Lee, Ching-Hua, et al.
Published: (2025)
by: Lee, Ching-Hua, et al.
Published: (2025)
Efficient Test-Time Adaptation through Latent Subspace Coefficients Search
by: Luo, Xinyu, et al.
Published: (2025)
by: Luo, Xinyu, et al.
Published: (2025)
A Smart-Glasses for Emergency Medical Services via Multimodal Multitask Learning
by: Jin, Liuyi, et al.
Published: (2025)
by: Jin, Liuyi, et al.
Published: (2025)
TACO: Rethinking Semantic Communications with Task Adaptation and Context Embedding
by: Wijesinghe, Achintha, et al.
Published: (2025)
by: Wijesinghe, Achintha, et al.
Published: (2025)
End-to-end audio-visual learning for cochlear implant sound coding simulations in noisy environments
by: Lin, Meng-Ping, et al.
Published: (2025)
by: Lin, Meng-Ping, et al.
Published: (2025)
Pitch-Conditioned Instrument Sound Synthesis From an Interactive Timbre Latent Space
by: Limberg, Christian, et al.
Published: (2025)
by: Limberg, Christian, et al.
Published: (2025)
Expressive Acoustic Guitar Sound Synthesis with an Instrument-Specific Input Representation and Diffusion Outpainting
by: Kim, Hounsu, et al.
Published: (2024)
by: Kim, Hounsu, et al.
Published: (2024)
Learning Video Temporal Dynamics with Cross-Modal Attention for Robust Audio-Visual Speech Recognition
by: Kim, Sungnyun, et al.
Published: (2024)
by: Kim, Sungnyun, et al.
Published: (2024)
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
by: Tuncay, Ludovic, et al.
Published: (2025)
by: Tuncay, Ludovic, et al.
Published: (2025)
When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds
by: Kang, Minsu, et al.
Published: (2025)
by: Kang, Minsu, et al.
Published: (2025)
Multimodal Marvels of Deep Learning in Medical Diagnosis: A Comprehensive Review of COVID-19 Detection
by: Islam, Md Shofiqul, et al.
Published: (2025)
by: Islam, Md Shofiqul, et al.
Published: (2025)
BUET Multi-disease Heart Sound Dataset: A Comprehensive Auscultation Dataset for Developing Computer-Aided Diagnostic Systems
by: Ali, Shams Nafisa, et al.
Published: (2024)
by: Ali, Shams Nafisa, et al.
Published: (2024)
Similar Items
-
Spatial and Semantic Embedding Integration for Stereo Sound Event Localization and Detection in Regular Videos
by: Berghi, Davide, et al.
Published: (2025) -
ToS: A Team of Specialists ensemble framework for Stereo Sound Event Localization and Detection with distance estimation in Video
by: Berghi, Davide, et al.
Published: (2026) -
Leveraging Reverberation and Visual Depth Cues for Sound Event Localization and Detection with Distance Estimation
by: Berghi, Davide, et al.
Published: (2024) -
Reverberation-based Features for Sound Event Localization and Detection with Distance Estimation
by: Berghi, Davide, et al.
Published: (2025) -
Quality-Controlled Multimodal Emotion Recognition in Conversations with Identity-Based Transfer Learning and MAMBA Fusion
by: Wang, Zanxu, et al.
Published: (2025)