AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jiayu, Ye, Shuo, Ye, Qilang, Lin, Xun, Song, Zihan, Yu, Zitong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Retrieving to Recover: Towards Incomplete Audio-Visual Question Answering via Semantic-consistent Purification
by: Zhang, Jiayu, et al.
Published: (2026)
by: Zhang, Jiayu, et al.
Published: (2026)
Answering Diverse Questions via Text Attached with Key Audio-Visual Clues
by: Ye, Qilang, et al.
Published: (2024)
by: Ye, Qilang, et al.
Published: (2024)
CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios
by: Ye, Qilang, et al.
Published: (2024)
by: Ye, Qilang, et al.
Published: (2024)
When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?
by: Ye, Qilang, et al.
Published: (2025)
by: Ye, Qilang, et al.
Published: (2025)
IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection
by: Li, Zewen, et al.
Published: (2025)
by: Li, Zewen, et al.
Published: (2025)
Question-Aware Gaussian Experts for Audio-Visual Question Answering
by: Kim, Hongyeob, et al.
Published: (2025)
by: Kim, Hongyeob, et al.
Published: (2025)
SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition
by: Ye, Qilang, et al.
Published: (2025)
by: Ye, Qilang, et al.
Published: (2025)
DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
by: Luo, Jingzhou, et al.
Published: (2025)
by: Luo, Jingzhou, et al.
Published: (2025)
A Comprehensive Survey on Visual Question Answering Datasets and Algorithms
by: Kabir, Raihan, et al.
Published: (2024)
by: Kabir, Raihan, et al.
Published: (2024)
Reconstruction as a Bridge for Event-Based Visual Question Answering
by: Lou, Hanyue, et al.
Published: (2025)
by: Lou, Hanyue, et al.
Published: (2025)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
by: Zhang, Chengyi, et al.
Published: (2026)
by: Zhang, Chengyi, et al.
Published: (2026)
InViC: Intent-aware Visual Cues for Medical Visual Question Answering
by: Wang, Zhisong, et al.
Published: (2026)
by: Wang, Zhisong, et al.
Published: (2026)
Multi-Sourced Compositional Generalization in Visual Question Answering
by: Li, Chuanhao, et al.
Published: (2025)
by: Li, Chuanhao, et al.
Published: (2025)
Efficient Visual Question Answering Pipeline for Autonomous Driving via Scene Region Compression
by: Cai, Yuliang, et al.
Published: (2026)
by: Cai, Yuliang, et al.
Published: (2026)
Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering
by: Ma, Jie, et al.
Published: (2024)
by: Ma, Jie, et al.
Published: (2024)
Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction Tuning
by: Zeng, Xingchen, et al.
Published: (2024)
by: Zeng, Xingchen, et al.
Published: (2024)
Patch-level Sounding Object Tracking for Audio-Visual Question Answering
by: Li, Zhangbin, et al.
Published: (2024)
by: Li, Zhangbin, et al.
Published: (2024)
EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning
by: Xing, Bohao, et al.
Published: (2024)
by: Xing, Bohao, et al.
Published: (2024)
MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
by: Song, Enxin, et al.
Published: (2024)
by: Song, Enxin, et al.
Published: (2024)
Visually Interpretable Subtask Reasoning for Visual Question Answering
by: Cheng, Yu, et al.
Published: (2025)
by: Cheng, Yu, et al.
Published: (2025)
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
by: Dang, Jisheng, et al.
Published: (2025)
by: Dang, Jisheng, et al.
Published: (2025)
DDAP: Dual-Domain Anti-Personalization against Text-to-Image Diffusion Models
by: Yang, Jing, et al.
Published: (2024)
by: Yang, Jing, et al.
Published: (2024)
YUV20K: A Complexity-Driven Benchmark and Trajectory-Aware Alignment Model for Video Camouflaged Object Detection
by: Liu, Yiyu, et al.
Published: (2026)
by: Liu, Yiyu, et al.
Published: (2026)
Query-Guided Spatial-Temporal-Frequency Interaction for Music Audio-Visual Question Answering
by: Li, Kun, et al.
Published: (2026)
by: Li, Kun, et al.
Published: (2026)
MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
by: Tang, Jingqun, et al.
Published: (2024)
by: Tang, Jingqun, et al.
Published: (2024)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
AutoVQA-G: Self-Improving Agentic Framework for Automated Visual Question Answering and Grounding Annotation
by: Hu, Rongsheng, et al.
Published: (2026)
by: Hu, Rongsheng, et al.
Published: (2026)
CLIP-Powered TASS: Target-Aware Single-Stream Network for Audio-Visual Question Answering
by: Jiang, Yuanyuan, et al.
Published: (2024)
by: Jiang, Yuanyuan, et al.
Published: (2024)
MedCFVQA: A Causal Approach to Mitigate Modality Preference Bias in Medical Visual Question Answering
by: Ye, Shuchang, et al.
Published: (2025)
by: Ye, Shuchang, et al.
Published: (2025)
AV-Flow: Transforming Text to Audio-Visual Human-like Interactions
by: Chatziagapi, Aggelina, et al.
Published: (2025)
by: Chatziagapi, Aggelina, et al.
Published: (2025)
VQAttack: Transferable Adversarial Attacks on Visual Question Answering via Pre-trained Models
by: Yin, Ziyi, et al.
Published: (2024)
by: Yin, Ziyi, et al.
Published: (2024)
Text-Guided Multimodal Unified Industrial Anomaly Detection
by: Li, Zewen, et al.
Published: (2026)
by: Li, Zewen, et al.
Published: (2026)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
by: Park, Kyu Ri, et al.
Published: (2024)
by: Park, Kyu Ri, et al.
Published: (2024)
AV-RIR: Audio-Visual Room Impulse Response Estimation
by: Ratnarajah, Anton, et al.
Published: (2023)
by: Ratnarajah, Anton, et al.
Published: (2023)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
by: Zhang, Xiaoman, et al.
Published: (2023)
by: Zhang, Xiaoman, et al.
Published: (2023)
PA-FAS: Towards Interpretable and Generalizable Multimodal Face Anti-Spoofing via Path-Augmented Reinforcement Learning
by: Ma, Yingjie, et al.
Published: (2025)
by: Ma, Yingjie, et al.
Published: (2025)
KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering
by: Li, Zhiyang, et al.
Published: (2026)
by: Li, Zhiyang, et al.
Published: (2026)
SaSR-Net: Source-Aware Semantic Representation Network for Enhancing Audio-Visual Question Answering
by: Yang, Tianyu, et al.
Published: (2024)
by: Yang, Tianyu, et al.
Published: (2024)
Questioning the Stability of Visual Question Answering
by: Rosenfeld, Amir, et al.
Published: (2025)
by: Rosenfeld, Amir, et al.
Published: (2025)
MedThink: Explaining Medical Visual Question Answering via Multimodal Decision-Making Rationale
by: Gai, Xiaotang, et al.
Published: (2024)
by: Gai, Xiaotang, et al.
Published: (2024)
Similar Items
-
Retrieving to Recover: Towards Incomplete Audio-Visual Question Answering via Semantic-consistent Purification
by: Zhang, Jiayu, et al.
Published: (2026) -
Answering Diverse Questions via Text Attached with Key Audio-Visual Clues
by: Ye, Qilang, et al.
Published: (2024) -
CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios
by: Ye, Qilang, et al.
Published: (2024) -
When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?
by: Ye, Qilang, et al.
Published: (2025) -
IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection
by: Li, Zewen, et al.
Published: (2025)