Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Chen, Li, Peike, Yang, Liying, Wang, Dadong, Li, Lincheng, Yu, Xin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915200645464064
author Liu, Chen
Li, Peike
Yang, Liying
Wang, Dadong
Li, Lincheng
Yu, Xin
author_facet Liu, Chen
Li, Peike
Yang, Liying
Wang, Dadong
Li, Lincheng
Yu, Xin
contents Accurately localizing audible objects based on audio-visual cues is the core objective of audio-visual segmentation. Most previous methods emphasize spatial or temporal multi-modal modeling, yet overlook challenges from ambiguous audio-visual correspondences such as nearby visually similar but acoustically different objects and frequent shifts in objects' sounding status. Consequently, they may struggle to reliably correlate audio and visual cues, leading to over- or under-segmentation. To address these limitations, we propose a novel framework with two primary components: an audio-guided modality alignment (AMA) module and an uncertainty estimation (UE) module. Instead of indiscriminately correlating audio-visual cues through a global attention mechanism, AMA performs audio-visual interactions within multiple groups and consolidates group features into compact representations based on their responsiveness to audio cues, effectively directing the model's attention to audio-relevant areas. Leveraging contrastive learning, AMA further distinguishes sounding regions from silent areas by treating features with strong audio responses as positive samples and weaker responses as negatives. Additionally, UE integrates spatial and temporal information to identify high-uncertainty regions caused by frequent changes in sound state, reducing prediction errors by lowering confidence in these areas. Experimental results demonstrate that our approach achieves superior accuracy compared to existing state-of-the-art methods, particularly in challenging scenarios where traditional approaches struggle to maintain reliable segmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12847
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment
Liu, Chen
Li, Peike
Yang, Liying
Wang, Dadong
Li, Lincheng
Yu, Xin
Sound
Computer Vision and Pattern Recognition
Accurately localizing audible objects based on audio-visual cues is the core objective of audio-visual segmentation. Most previous methods emphasize spatial or temporal multi-modal modeling, yet overlook challenges from ambiguous audio-visual correspondences such as nearby visually similar but acoustically different objects and frequent shifts in objects' sounding status. Consequently, they may struggle to reliably correlate audio and visual cues, leading to over- or under-segmentation. To address these limitations, we propose a novel framework with two primary components: an audio-guided modality alignment (AMA) module and an uncertainty estimation (UE) module. Instead of indiscriminately correlating audio-visual cues through a global attention mechanism, AMA performs audio-visual interactions within multiple groups and consolidates group features into compact representations based on their responsiveness to audio cues, effectively directing the model's attention to audio-relevant areas. Leveraging contrastive learning, AMA further distinguishes sounding regions from silent areas by treating features with strong audio responses as positive samples and weaker responses as negatives. Additionally, UE integrates spatial and temporal information to identify high-uncertainty regions caused by frequent changes in sound state, reducing prediction errors by lowering confidence in these areas. Experimental results demonstrate that our approach achieves superior accuracy compared to existing state-of-the-art methods, particularly in challenging scenarios where traditional approaches struggle to maintain reliable segmentation.
title Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment
topic Sound
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.12847