Object-aware Sound Source Localization via Audio-Visual Scene Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Um, Sung Jin, Kim, Dongjin, Lee, Sangmin, Kim, Jung Uk
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916808888418304
author Um, Sung Jin
Kim, Dongjin
Lee, Sangmin
Kim, Jung Uk
author_facet Um, Sung Jin
Kim, Dongjin
Lee, Sangmin
Kim, Jung Uk
contents Audio-visual sound source localization task aims to spatially localize sound-making objects within visual scenes by integrating visual and audio cues. However, existing methods struggle with accurately localizing sound-making objects in complex scenes, particularly when visually similar silent objects coexist. This limitation arises primarily from their reliance on simple audio-visual correspondence, which does not capture fine-grained semantic differences between sound-making and silent objects. To address these challenges, we propose a novel sound source localization framework leveraging Multimodal Large Language Models (MLLMs) to generate detailed contextual information that explicitly distinguishes between sound-making foreground objects and silent background objects. To effectively integrate this detailed information, we introduce two novel loss functions: Object-aware Contrastive Alignment (OCA) loss and Object Region Isolation (ORI) loss. Extensive experimental results on MUSIC and VGGSound datasets demonstrate the effectiveness of our approach, significantly outperforming existing methods in both single-source and multi-source localization scenarios. Code and generated detailed contextual information are available at: https://github.com/VisualAIKHU/OA-SSL.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18557
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Object-aware Sound Source Localization via Audio-Visual Scene Understanding
Um, Sung Jin
Kim, Dongjin
Lee, Sangmin
Kim, Jung Uk
Computer Vision and Pattern Recognition
Audio-visual sound source localization task aims to spatially localize sound-making objects within visual scenes by integrating visual and audio cues. However, existing methods struggle with accurately localizing sound-making objects in complex scenes, particularly when visually similar silent objects coexist. This limitation arises primarily from their reliance on simple audio-visual correspondence, which does not capture fine-grained semantic differences between sound-making and silent objects. To address these challenges, we propose a novel sound source localization framework leveraging Multimodal Large Language Models (MLLMs) to generate detailed contextual information that explicitly distinguishes between sound-making foreground objects and silent background objects. To effectively integrate this detailed information, we introduce two novel loss functions: Object-aware Contrastive Alignment (OCA) loss and Object Region Isolation (ORI) loss. Extensive experimental results on MUSIC and VGGSound datasets demonstrate the effectiveness of our approach, significantly outperforming existing methods in both single-source and multi-source localization scenarios. Code and generated detailed contextual information are available at: https://github.com/VisualAIKHU/OA-SSL.
title Object-aware Sound Source Localization via Audio-Visual Scene Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.18557