AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Jiayu, Ye, Shuo, Ye, Qilang, Lin, Xun, Song, Zihan, Yu, Zitong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908992424378368
author Zhang, Jiayu
Ye, Shuo
Ye, Qilang
Lin, Xun
Song, Zihan
Yu, Zitong
author_facet Zhang, Jiayu
Ye, Shuo
Ye, Qilang
Lin, Xun
Song, Zihan
Yu, Zitong
contents Audio-Visual Question Answering (AVQA) requires models to effectively utilize both visual and auditory modalities to answer complex and diverse questions about audio-visual scenes. However, existing methods lack sufficient flexibility and dynamic adaptability in temporal sampling and modality preference awareness, making it difficult to focus on key information based on the question. This limits their reasoning capability in complex scenarios. To address these challenges, we propose a novel framework named AV-Master. It enhances the model's ability to extract key information from complex audio-visual scenes with substantial redundant content by dynamically modeling both temporal and modality dimensions. In the temporal dimension, we introduce a dynamic adaptive focus sampling mechanism that progressively focuses on audio-visual segments most relevant to the question, effectively mitigating redundancy and segment fragmentation in traditional sampling methods. In the modality dimension, we propose a preference-aware strategy that models each modality's contribution independently, enabling selective activation of critical features. Furthermore, we introduce a dual-path contrastive loss to reinforce consistency and complementarity across temporal and modality dimensions, guiding the model to learn question-specific cross-modal collaborative representations. Experiments on four large-scale benchmarks show that AV-Master significantly outperforms existing methods, especially in complex reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18346
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering
Zhang, Jiayu
Ye, Shuo
Ye, Qilang
Lin, Xun
Song, Zihan
Yu, Zitong
Computer Vision and Pattern Recognition
Audio-Visual Question Answering (AVQA) requires models to effectively utilize both visual and auditory modalities to answer complex and diverse questions about audio-visual scenes. However, existing methods lack sufficient flexibility and dynamic adaptability in temporal sampling and modality preference awareness, making it difficult to focus on key information based on the question. This limits their reasoning capability in complex scenarios. To address these challenges, we propose a novel framework named AV-Master. It enhances the model's ability to extract key information from complex audio-visual scenes with substantial redundant content by dynamically modeling both temporal and modality dimensions. In the temporal dimension, we introduce a dynamic adaptive focus sampling mechanism that progressively focuses on audio-visual segments most relevant to the question, effectively mitigating redundancy and segment fragmentation in traditional sampling methods. In the modality dimension, we propose a preference-aware strategy that models each modality's contribution independently, enabling selective activation of critical features. Furthermore, we introduce a dual-path contrastive loss to reinforce consistency and complementarity across temporal and modality dimensions, guiding the model to learn question-specific cross-modal collaborative representations. Experiments on four large-scale benchmarks show that AV-Master significantly outperforms existing methods, especially in complex reasoning tasks.
title AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.18346