Cooperation Does Matter: Exploring Multi-Order Bilateral Relations for Audio-Visual Segmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Qi, Nie, Xing, Li, Tong, Gao, Pengfei, Guo, Ying, Zhen, Cheng, Yan, Pengfei, Xiang, Shiming
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909162825318400
author Yang, Qi
Nie, Xing
Li, Tong
Gao, Pengfei
Guo, Ying
Zhen, Cheng
Yan, Pengfei
Xiang, Shiming
author_facet Yang, Qi
Nie, Xing
Li, Tong
Gao, Pengfei
Guo, Ying
Zhen, Cheng
Yan, Pengfei
Xiang, Shiming
contents Recently, an audio-visual segmentation (AVS) task has been introduced, aiming to group pixels with sounding objects within a given video. This task necessitates a first-ever audio-driven pixel-level understanding of the scene, posing significant challenges. In this paper, we propose an innovative audio-visual transformer framework, termed COMBO, an acronym for COoperation of Multi-order Bilateral relatiOns. For the first time, our framework explores three types of bilateral entanglements within AVS: pixel entanglement, modality entanglement, and temporal entanglement. Regarding pixel entanglement, we employ a Siam-Encoder Module (SEM) that leverages prior knowledge to generate more precise visual features from the foundational model. For modality entanglement, we design a Bilateral-Fusion Module (BFM), enabling COMBO to align corresponding visual and auditory signals bi-directionally. As for temporal entanglement, we introduce an innovative adaptive inter-frame consistency loss according to the inherent rules of temporal. Comprehensive experiments and ablation studies on AVSBench-object (84.7 mIoU on S4, 59.2 mIou on MS3) and AVSBench-semantic (42.1 mIoU on AVSS) datasets demonstrate that COMBO surpasses previous state-of-the-art methods. Code and more results will be publicly available at https://yannqi.github.io/AVS-COMBO/.
format Preprint
id arxiv_https___arxiv_org_abs_2312_06462
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Cooperation Does Matter: Exploring Multi-Order Bilateral Relations for Audio-Visual Segmentation
Yang, Qi
Nie, Xing
Li, Tong
Gao, Pengfei
Guo, Ying
Zhen, Cheng
Yan, Pengfei
Xiang, Shiming
Computer Vision and Pattern Recognition
Artificial Intelligence
Sound
Audio and Speech Processing
Recently, an audio-visual segmentation (AVS) task has been introduced, aiming to group pixels with sounding objects within a given video. This task necessitates a first-ever audio-driven pixel-level understanding of the scene, posing significant challenges. In this paper, we propose an innovative audio-visual transformer framework, termed COMBO, an acronym for COoperation of Multi-order Bilateral relatiOns. For the first time, our framework explores three types of bilateral entanglements within AVS: pixel entanglement, modality entanglement, and temporal entanglement. Regarding pixel entanglement, we employ a Siam-Encoder Module (SEM) that leverages prior knowledge to generate more precise visual features from the foundational model. For modality entanglement, we design a Bilateral-Fusion Module (BFM), enabling COMBO to align corresponding visual and auditory signals bi-directionally. As for temporal entanglement, we introduce an innovative adaptive inter-frame consistency loss according to the inherent rules of temporal. Comprehensive experiments and ablation studies on AVSBench-object (84.7 mIoU on S4, 59.2 mIou on MS3) and AVSBench-semantic (42.1 mIoU on AVSS) datasets demonstrate that COMBO surpasses previous state-of-the-art methods. Code and more results will be publicly available at https://yannqi.github.io/AVS-COMBO/.
title Cooperation Does Matter: Exploring Multi-Order Bilateral Relations for Audio-Visual Segmentation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2312.06462