Frequency-Domain Decomposition and Recomposition for Robust Audio-Visual Segmentation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shen, Yunzhe, Peng, Kai, Liu, Leiye, Ji, Wei, Li, Jingjing, Zhang, Miao, Piao, Yongri, Lu, Huchuan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918146412118016
author Shen, Yunzhe
Peng, Kai
Liu, Leiye
Ji, Wei
Li, Jingjing
Zhang, Miao
Piao, Yongri
Lu, Huchuan
author_facet Shen, Yunzhe
Peng, Kai
Liu, Leiye
Ji, Wei
Li, Jingjing
Zhang, Miao
Piao, Yongri
Lu, Huchuan
contents Audio-visual segmentation (AVS) plays a critical role in multimodal machine learning by effectively integrating audio and visual cues to precisely segment objects or regions within visual scenes. Recent AVS methods have demonstrated significant improvements. However, they overlook the inherent frequency-domain contradictions between audio and visual modalities--the pervasively interfering noise in audio high-frequency signals vs. the structurally rich details in visual high-frequency signals. Ignoring these differences can result in suboptimal performance. In this paper, we rethink the AVS task from a deeper perspective by reformulating AVS task as a frequency-domain decomposition and recomposition problem. To this end, we introduce a novel Frequency-Aware Audio-Visual Segmentation (FAVS) framework consisting of two key modules: Frequency-Domain Enhanced Decomposer (FDED) module and Synergistic Cross-Modal Consistency (SCMC) module. FDED module employs a residual-based iterative frequency decomposition to discriminate modality-specific semantics and structural features, and SCMC module leverages a mixture-of-experts architecture to reinforce semantic consistency and modality-specific feature preservation through dynamic expert routing. Extensive experiments demonstrate that our FAVS framework achieves state-of-the-art performance on three benchmark datasets, and abundant qualitative visualizations further verify the effectiveness of the proposed FDED and SCMC modules. The code will be released as open source upon acceptance of the paper.
format Preprint
id arxiv_https___arxiv_org_abs_2509_18912
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Frequency-Domain Decomposition and Recomposition for Robust Audio-Visual Segmentation
Shen, Yunzhe
Peng, Kai
Liu, Leiye
Ji, Wei
Li, Jingjing
Zhang, Miao
Piao, Yongri
Lu, Huchuan
Computer Vision and Pattern Recognition
Audio-visual segmentation (AVS) plays a critical role in multimodal machine learning by effectively integrating audio and visual cues to precisely segment objects or regions within visual scenes. Recent AVS methods have demonstrated significant improvements. However, they overlook the inherent frequency-domain contradictions between audio and visual modalities--the pervasively interfering noise in audio high-frequency signals vs. the structurally rich details in visual high-frequency signals. Ignoring these differences can result in suboptimal performance. In this paper, we rethink the AVS task from a deeper perspective by reformulating AVS task as a frequency-domain decomposition and recomposition problem. To this end, we introduce a novel Frequency-Aware Audio-Visual Segmentation (FAVS) framework consisting of two key modules: Frequency-Domain Enhanced Decomposer (FDED) module and Synergistic Cross-Modal Consistency (SCMC) module. FDED module employs a residual-based iterative frequency decomposition to discriminate modality-specific semantics and structural features, and SCMC module leverages a mixture-of-experts architecture to reinforce semantic consistency and modality-specific feature preservation through dynamic expert routing. Extensive experiments demonstrate that our FAVS framework achieves state-of-the-art performance on three benchmark datasets, and abundant qualitative visualizations further verify the effectiveness of the proposed FDED and SCMC modules. The code will be released as open source upon acceptance of the paper.
title Frequency-Domain Decomposition and Recomposition for Robust Audio-Visual Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.18912