DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xiong, Junwen, Zhang, Peng, You, Tao, Li, Chuanyue, Huang, Wei, Zha, Yufei
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913251674030080
author Xiong, Junwen
Zhang, Peng
You, Tao
Li, Chuanyue
Huang, Wei
Zha, Yufei
author_facet Xiong, Junwen
Zhang, Peng
You, Tao
Li, Chuanyue
Huang, Wei
Zha, Yufei
contents Audio-visual saliency prediction can draw support from diverse modality complements, but further performance enhancement is still challenged by customized architectures as well as task-specific loss functions. In recent studies, denoising diffusion models have shown more promising in unifying task frameworks owing to their inherent ability of generalization. Following this motivation, a novel Diffusion architecture for generalized audio-visual Saliency prediction (DiffSal) is proposed in this work, which formulates the prediction problem as a conditional generative task of the saliency map by utilizing input audio and video as the conditions. Based on the spatio-temporal audio-visual features, an extra network Saliency-UNet is designed to perform multi-modal attention modulation for progressive refinement of the ground-truth saliency map from the noisy map. Extensive experiments demonstrate that the proposed DiffSal can achieve excellent performance across six challenging audio-visual benchmarks, with an average relative improvement of 6.3\% over the previous state-of-the-art results by six metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2403_01226
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction
Xiong, Junwen
Zhang, Peng
You, Tao
Li, Chuanyue
Huang, Wei
Zha, Yufei
Computer Vision and Pattern Recognition
Audio-visual saliency prediction can draw support from diverse modality complements, but further performance enhancement is still challenged by customized architectures as well as task-specific loss functions. In recent studies, denoising diffusion models have shown more promising in unifying task frameworks owing to their inherent ability of generalization. Following this motivation, a novel Diffusion architecture for generalized audio-visual Saliency prediction (DiffSal) is proposed in this work, which formulates the prediction problem as a conditional generative task of the saliency map by utilizing input audio and video as the conditions. Based on the spatio-temporal audio-visual features, an extra network Saliency-UNet is designed to perform multi-modal attention modulation for progressive refinement of the ground-truth saliency map from the noisy map. Extensive experiments demonstrate that the proposed DiffSal can achieve excellent performance across six challenging audio-visual benchmarks, with an average relative improvement of 6.3\% over the previous state-of-the-art results by six metrics.
title DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.01226