Extending Segment Anything Model into Auditory and Temporal Dimensions for Audio-Visual Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Seon, Juhyeong, Im, Woobin, Lee, Sebin, Lee, Jumin, Yoon, Sung-Eui
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929380326899712
author Seon, Juhyeong
Im, Woobin
Lee, Sebin
Lee, Jumin
Yoon, Sung-Eui
author_facet Seon, Juhyeong
Im, Woobin
Lee, Sebin
Lee, Jumin
Yoon, Sung-Eui
contents Audio-visual segmentation (AVS) aims to segment sound sources in the video sequence, requiring a pixel-level understanding of audio-visual correspondence. As the Segment Anything Model (SAM) has strongly impacted extensive fields of dense prediction problems, prior works have investigated the introduction of SAM into AVS with audio as a new modality of the prompt. Nevertheless, constrained by SAM's single-frame segmentation scheme, the temporal context across multiple frames of audio-visual data remains insufficiently utilized. To this end, we study the extension of SAM's capabilities to the sequence of audio-visual scenes by analyzing contextual cross-modal relationships across the frames. To achieve this, we propose a Spatio-Temporal, Bidirectional Audio-Visual Attention (ST-BAVA) module integrated into the middle of SAM's image encoder and mask decoder. It adaptively updates the audio-visual features to convey the spatio-temporal correspondence between the video frames and audio streams. Extensive experiments demonstrate that our proposed model outperforms the state-of-the-art methods on AVS benchmarks, especially with an 8.3% mIoU gain on a challenging multi-sources subset.
format Preprint
id arxiv_https___arxiv_org_abs_2406_06163
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Extending Segment Anything Model into Auditory and Temporal Dimensions for Audio-Visual Segmentation
Seon, Juhyeong
Im, Woobin
Lee, Sebin
Lee, Jumin
Yoon, Sung-Eui
Computer Vision and Pattern Recognition
Audio-visual segmentation (AVS) aims to segment sound sources in the video sequence, requiring a pixel-level understanding of audio-visual correspondence. As the Segment Anything Model (SAM) has strongly impacted extensive fields of dense prediction problems, prior works have investigated the introduction of SAM into AVS with audio as a new modality of the prompt. Nevertheless, constrained by SAM's single-frame segmentation scheme, the temporal context across multiple frames of audio-visual data remains insufficiently utilized. To this end, we study the extension of SAM's capabilities to the sequence of audio-visual scenes by analyzing contextual cross-modal relationships across the frames. To achieve this, we propose a Spatio-Temporal, Bidirectional Audio-Visual Attention (ST-BAVA) module integrated into the middle of SAM's image encoder and mask decoder. It adaptively updates the audio-visual features to convey the spatio-temporal correspondence between the video frames and audio streams. Extensive experiments demonstrate that our proposed model outperforms the state-of-the-art methods on AVS benchmarks, especially with an 8.3% mIoU gain on a challenging multi-sources subset.
title Extending Segment Anything Model into Auditory and Temporal Dimensions for Audio-Visual Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.06163