Progressive Confident Masking Attention Network for Audio-Visual Segmentation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Yuxuan, Zhu, Jinchao, Dong, Feng, Zhu, Shuyue
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916605289562112
author Wang, Yuxuan
Zhu, Jinchao
Dong, Feng
Zhu, Shuyue
author_facet Wang, Yuxuan
Zhu, Jinchao
Dong, Feng
Zhu, Shuyue
contents Audio and visual signals typically occur simultaneously, and humans possess an innate ability to correlate and synchronize information from these two modalities. Recently, a challenging problem known as Audio-Visual Segmentation (AVS) has emerged, intending to produce segmentation maps for sounding objects within a scene. However, the methods proposed so far have not sufficiently integrated audio and visual information, and the computational costs have been extremely high. Additionally, the outputs of different stages have not been fully utilized. To facilitate this research, we introduce a novel Progressive Confident Masking Attention Network (PMCANet). It leverages attention mechanisms to uncover the intrinsic correlations between audio signals and visual frames. Furthermore, we design an efficient and effective cross-attention module to enhance semantic perception by selecting query tokens. This selection is determined through confidence-driven units based on the network's multi-stage predictive outputs. Experiments demonstrate that our network outperforms other AVS methods while requiring less computational resources. The code is available at: https://github.com/PrettyPlate/PCMANet.
format Preprint
id arxiv_https___arxiv_org_abs_2406_02345
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Progressive Confident Masking Attention Network for Audio-Visual Segmentation
Wang, Yuxuan
Zhu, Jinchao
Dong, Feng
Zhu, Shuyue
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimedia
Audio and visual signals typically occur simultaneously, and humans possess an innate ability to correlate and synchronize information from these two modalities. Recently, a challenging problem known as Audio-Visual Segmentation (AVS) has emerged, intending to produce segmentation maps for sounding objects within a scene. However, the methods proposed so far have not sufficiently integrated audio and visual information, and the computational costs have been extremely high. Additionally, the outputs of different stages have not been fully utilized. To facilitate this research, we introduce a novel Progressive Confident Masking Attention Network (PMCANet). It leverages attention mechanisms to uncover the intrinsic correlations between audio signals and visual frames. Furthermore, we design an efficient and effective cross-attention module to enhance semantic perception by selecting query tokens. This selection is determined through confidence-driven units based on the network's multi-stage predictive outputs. Experiments demonstrate that our network outperforms other AVS methods while requiring less computational resources. The code is available at: https://github.com/PrettyPlate/PCMANet.
title Progressive Confident Masking Attention Network for Audio-Visual Segmentation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimedia
url https://arxiv.org/abs/2406.02345