SaSR-Net: Source-Aware Semantic Representation Network for Enhancing Audio-Visual Question Answering

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Tianyu, Nan, Yiyang, Dai, Lisen, Liang, Zhenwen, Tian, Yapeng, Zhang, Xiangliang
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929585527980032
author Yang, Tianyu
Nan, Yiyang
Dai, Lisen
Liang, Zhenwen
Tian, Yapeng
Zhang, Xiangliang
author_facet Yang, Tianyu
Nan, Yiyang
Dai, Lisen
Liang, Zhenwen
Tian, Yapeng
Zhang, Xiangliang
contents Audio-Visual Question Answering (AVQA) is a challenging task that involves answering questions based on both auditory and visual information in videos. A significant challenge is interpreting complex multi-modal scenes, which include both visual objects and sound sources, and connecting them to the given question. In this paper, we introduce the Source-aware Semantic Representation Network (SaSR-Net), a novel model designed for AVQA. SaSR-Net utilizes source-wise learnable tokens to efficiently capture and align audio-visual elements with the corresponding question. It streamlines the fusion of audio and visual information using spatial and temporal attention mechanisms to identify answers in multi-modal scenes. Extensive experiments on the Music-AVQA and AVQA-Yang datasets show that SaSR-Net outperforms state-of-the-art AVQA methods.
format Preprint
id arxiv_https___arxiv_org_abs_2411_04933
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SaSR-Net: Source-Aware Semantic Representation Network for Enhancing Audio-Visual Question Answering
Yang, Tianyu
Nan, Yiyang
Dai, Lisen
Liang, Zhenwen
Tian, Yapeng
Zhang, Xiangliang
Computer Vision and Pattern Recognition
Audio-Visual Question Answering (AVQA) is a challenging task that involves answering questions based on both auditory and visual information in videos. A significant challenge is interpreting complex multi-modal scenes, which include both visual objects and sound sources, and connecting them to the given question. In this paper, we introduce the Source-aware Semantic Representation Network (SaSR-Net), a novel model designed for AVQA. SaSR-Net utilizes source-wise learnable tokens to efficiently capture and align audio-visual elements with the corresponding question. It streamlines the fusion of audio and visual information using spatial and temporal attention mechanisms to identify answers in multi-modal scenes. Extensive experiments on the Music-AVQA and AVQA-Yang datasets show that SaSR-Net outperforms state-of-the-art AVQA methods.
title SaSR-Net: Source-Aware Semantic Representation Network for Enhancing Audio-Visual Question Answering
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.04933