A Critical Assessment of Visual Sound Source Localization Models Including Negative Audio
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Juanola, Xavier, Haro, Gloria, Fuentes, Magdalena |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment
von: Senocak, Arda, et al.
Veröffentlicht: (2024)
von: Senocak, Arda, et al.
Veröffentlicht: (2024)
Audio-Visual Talker Localization in Video for Spatial Sound Reproduction
von: Berghi, Davide, et al.
Veröffentlicht: (2024)
von: Berghi, Davide, et al.
Veröffentlicht: (2024)
T-VSL: Text-Guided Visual Sound Source Localization in Mixtures
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer
von: Wang, Yaoting, et al.
Veröffentlicht: (2023)
von: Wang, Yaoting, et al.
Veröffentlicht: (2023)
Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes
von: Ryu, Hyeonggon, et al.
Veröffentlicht: (2025)
von: Ryu, Hyeonggon, et al.
Veröffentlicht: (2025)
Learning to Visually Localize Sound Sources from Mixtures without Prior Source Knowledge
von: Kim, Dongjin, et al.
Veröffentlicht: (2024)
von: Kim, Dongjin, et al.
Veröffentlicht: (2024)
Cross Pseudo-Labeling for Semi-Supervised Audio-Visual Source Localization
von: Guo, Yuxin, et al.
Veröffentlicht: (2024)
von: Guo, Yuxin, et al.
Veröffentlicht: (2024)
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation
von: Petermann, Darius, et al.
Veröffentlicht: (2025)
von: Petermann, Darius, et al.
Veröffentlicht: (2025)
Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization
von: Park, Sooyoung, et al.
Veröffentlicht: (2025)
von: Park, Sooyoung, et al.
Veröffentlicht: (2025)
Learning Self-Supervised Audio-Visual Representations for Sound Recommendations
von: Krishnamurthy, Sudha
Veröffentlicht: (2024)
von: Krishnamurthy, Sudha
Veröffentlicht: (2024)
Pindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine Grained-Localization
von: Klein, Nicholas, et al.
Veröffentlicht: (2025)
von: Klein, Nicholas, et al.
Veröffentlicht: (2025)
Gotta Hear Them All: Towards Sound Source Aware Audio Generation
von: Guo, Wei, et al.
Veröffentlicht: (2024)
von: Guo, Wei, et al.
Veröffentlicht: (2024)
Global-Local Distillation Network-Based Audio-Visual Speaker Tracking with Incomplete Modalities
von: Li, Yidi, et al.
Veröffentlicht: (2024)
von: Li, Yidi, et al.
Veröffentlicht: (2024)
Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
von: Liu, Huadai, et al.
Veröffentlicht: (2025)
von: Liu, Huadai, et al.
Veröffentlicht: (2025)
Segmenting Collision Sound Sources in Egocentric Videos
von: Parida, Kranti Kumar, et al.
Veröffentlicht: (2025)
von: Parida, Kranti Kumar, et al.
Veröffentlicht: (2025)
SoundWeaver: Semantic Warm-Starting for Text-to-Audio Diffusion Serving
von: Barik, Ayush, et al.
Veröffentlicht: (2026)
von: Barik, Ayush, et al.
Veröffentlicht: (2026)
CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation
von: Chen, Yuanhong, et al.
Veröffentlicht: (2025)
von: Chen, Yuanhong, et al.
Veröffentlicht: (2025)
Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio Semantics
von: Liu, Chen, et al.
Veröffentlicht: (2025)
von: Liu, Chen, et al.
Veröffentlicht: (2025)
SemiPL: A Semi-supervised Method for Event Sound Source Localization
von: Li, Yue, et al.
Veröffentlicht: (2024)
von: Li, Yue, et al.
Veröffentlicht: (2024)
DDAVS: Disentangled Audio Semantics and Delayed Bidirectional Alignment for Audio-Visual Segmentation
von: Tian, Jingqi, et al.
Veröffentlicht: (2025)
von: Tian, Jingqi, et al.
Veröffentlicht: (2025)
SSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification
von: Rajasekhar, Gnana Praveen, et al.
Veröffentlicht: (2025)
von: Rajasekhar, Gnana Praveen, et al.
Veröffentlicht: (2025)
AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
UniSync: A Unified Framework for Audio-Visual Synchronization
von: Feng, Tao, et al.
Veröffentlicht: (2025)
von: Feng, Tao, et al.
Veröffentlicht: (2025)
Measuring Sound Symbolism in Audio-visual Models
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2024)
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2024)
RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation
von: Pegg, Samuel, et al.
Veröffentlicht: (2023)
von: Pegg, Samuel, et al.
Veröffentlicht: (2023)
Pilot-guided Multimodal Semantic Communication for Audio-Visual Event Localization
von: Yu, Fei, et al.
Veröffentlicht: (2024)
von: Yu, Fei, et al.
Veröffentlicht: (2024)
Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2024)
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2024)
DeepSound-V1: Start to Think Step-by-Step in the Audio Generation from Videos
von: Liang, Yunming, et al.
Veröffentlicht: (2025)
von: Liang, Yunming, et al.
Veröffentlicht: (2025)
Localizing Audio-Visual Deepfakes via Hierarchical Boundary Modeling
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
What's Making That Sound Right Now? Video-centric Audio-Visual Localization
von: Choi, Hahyeon, et al.
Veröffentlicht: (2025)
von: Choi, Hahyeon, et al.
Veröffentlicht: (2025)
Look, Listen and Recognise: Character-Aware Audio-Visual Subtitling
von: Korbar, Bruno, et al.
Veröffentlicht: (2024)
von: Korbar, Bruno, et al.
Veröffentlicht: (2024)
Continual Audio-Visual Sound Separation
von: Pian, Weiguo, et al.
Veröffentlicht: (2024)
von: Pian, Weiguo, et al.
Veröffentlicht: (2024)
UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing
von: Lai, Yung-Hsuan, et al.
Veröffentlicht: (2025)
von: Lai, Yung-Hsuan, et al.
Veröffentlicht: (2025)
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos
von: Majumder, Sagnik, et al.
Veröffentlicht: (2023)
von: Majumder, Sagnik, et al.
Veröffentlicht: (2023)
SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
von: He, Yuhang, et al.
Veröffentlicht: (2024)
von: He, Yuhang, et al.
Veröffentlicht: (2024)
High-Quality Visually-Guided Sound Separation from Diverse Categories
von: Huang, Chao, et al.
Veröffentlicht: (2023)
von: Huang, Chao, et al.
Veröffentlicht: (2023)
UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
von: Geng, Tiantian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment
von: Senocak, Arda, et al.
Veröffentlicht: (2024) -
Audio-Visual Talker Localization in Video for Spatial Sound Reproduction
von: Berghi, Davide, et al.
Veröffentlicht: (2024) -
T-VSL: Text-Guided Visual Sound Source Localization in Mixtures
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024) -
Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer
von: Wang, Yaoting, et al.
Veröffentlicht: (2023) -
Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes
von: Ryu, Hyeonggon, et al.
Veröffentlicht: (2025)