Selective Noise Suppression and Discriminative Mutual Interaction for Robust Audio-Visual Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Kai, Shen, Yunzhe, Zhang, Miao, Liu, Leiye, Han, Yidong, Ji, Wei, Li, Jingjing, Piao, Yongri, Lu, Huchuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915885587890176
author Peng, Kai
Shen, Yunzhe
Zhang, Miao
Liu, Leiye
Han, Yidong
Ji, Wei
Li, Jingjing
Piao, Yongri
Lu, Huchuan
author_facet Peng, Kai
Shen, Yunzhe
Zhang, Miao
Liu, Leiye
Han, Yidong
Ji, Wei
Li, Jingjing
Piao, Yongri
Lu, Huchuan
contents The ability to capture and segment sounding objects in dynamic visual scenes is crucial for the development of Audio-Visual Segmentation (AVS) tasks. While significant progress has been made in this area, the interaction between audio and visual modalities still requires further exploration. In this work, we aim to answer the following questions: How can a model effectively suppress audio noise while enhancing relevant audio information? How can we achieve discriminative interaction between the audio and visual modalities? To this end, we propose SDAVS, equipped with the Selective Noise-Resilient Processor (SNRP) module and the Discriminative Audio-Visual Mutual Fusion (DAMF) strategy. The proposed SNRP mitigates audio noise interference by selectively emphasizing relevant auditory cues, while DAMF ensures more consistent audio-visual representations. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on benchmark AVS datasets, especially in multi-source and complex scenes. \textit{The code and model are available at https://github.com/happylife-pk/SDAVS}.
format Preprint
id arxiv_https___arxiv_org_abs_2603_14203
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Selective Noise Suppression and Discriminative Mutual Interaction for Robust Audio-Visual Segmentation
Peng, Kai
Shen, Yunzhe
Zhang, Miao
Liu, Leiye
Han, Yidong
Ji, Wei
Li, Jingjing
Piao, Yongri
Lu, Huchuan
Computer Vision and Pattern Recognition
The ability to capture and segment sounding objects in dynamic visual scenes is crucial for the development of Audio-Visual Segmentation (AVS) tasks. While significant progress has been made in this area, the interaction between audio and visual modalities still requires further exploration. In this work, we aim to answer the following questions: How can a model effectively suppress audio noise while enhancing relevant audio information? How can we achieve discriminative interaction between the audio and visual modalities? To this end, we propose SDAVS, equipped with the Selective Noise-Resilient Processor (SNRP) module and the Discriminative Audio-Visual Mutual Fusion (DAMF) strategy. The proposed SNRP mitigates audio noise interference by selectively emphasizing relevant auditory cues, while DAMF ensures more consistent audio-visual representations. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on benchmark AVS datasets, especially in multi-source and complex scenes. \textit{The code and model are available at https://github.com/happylife-pk/SDAVS}.
title Selective Noise Suppression and Discriminative Mutual Interaction for Robust Audio-Visual Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.14203