Revisiting Audio-Visual Segmentation with Vision-Centric Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Shaofei, Ling, Rui, Hui, Tianrui, Li, Hongyu, Zhou, Xu, Zhang, Shifeng, Liu, Si, Hong, Richang, Wang, Meng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909666963881984
author Huang, Shaofei
Ling, Rui
Hui, Tianrui
Li, Hongyu
Zhou, Xu
Zhang, Shifeng
Liu, Si
Hong, Richang
Wang, Meng
author_facet Huang, Shaofei
Ling, Rui
Hui, Tianrui
Li, Hongyu
Zhou, Xu
Zhang, Shifeng
Liu, Si
Hong, Richang
Wang, Meng
contents Audio-Visual Segmentation (AVS) aims to segment sound-producing objects in video frames based on the associated audio signal. Prevailing AVS methods typically adopt an audio-centric Transformer architecture, where object queries are derived from audio features. However, audio-centric Transformers suffer from two limitations: perception ambiguity caused by the mixed nature of audio, and weakened dense prediction ability due to visual detail loss. To address these limitations, we propose a new Vision-Centric Transformer (VCT) framework that leverages vision-derived queries to iteratively fetch corresponding audio and visual information, enabling queries to better distinguish between different sounding objects from mixed audio and accurately delineate their contours. Additionally, we also introduce a Prototype Prompted Query Generation (PPQG) module within our VCT framework to generate vision-derived queries that are both semantically aware and visually rich through audio prototype prompting and pixel context grouping, facilitating audio-visual information aggregation. Extensive experiments demonstrate that our VCT framework achieves new state-of-the-art performances on three subsets of the AVSBench dataset. The code is available at https://github.com/spyflying/VCT_AVS.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23623
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revisiting Audio-Visual Segmentation with Vision-Centric Transformer
Huang, Shaofei
Ling, Rui
Hui, Tianrui
Li, Hongyu
Zhou, Xu
Zhang, Shifeng
Liu, Si
Hong, Richang
Wang, Meng
Computer Vision and Pattern Recognition
Audio-Visual Segmentation (AVS) aims to segment sound-producing objects in video frames based on the associated audio signal. Prevailing AVS methods typically adopt an audio-centric Transformer architecture, where object queries are derived from audio features. However, audio-centric Transformers suffer from two limitations: perception ambiguity caused by the mixed nature of audio, and weakened dense prediction ability due to visual detail loss. To address these limitations, we propose a new Vision-Centric Transformer (VCT) framework that leverages vision-derived queries to iteratively fetch corresponding audio and visual information, enabling queries to better distinguish between different sounding objects from mixed audio and accurately delineate their contours. Additionally, we also introduce a Prototype Prompted Query Generation (PPQG) module within our VCT framework to generate vision-derived queries that are both semantically aware and visually rich through audio prototype prompting and pixel context grouping, facilitating audio-visual information aggregation. Extensive experiments demonstrate that our VCT framework achieves new state-of-the-art performances on three subsets of the AVSBench dataset. The code is available at https://github.com/spyflying/VCT_AVS.
title Revisiting Audio-Visual Segmentation with Vision-Centric Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.23623