Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering
Fuente:
arXiv
Salvato in:
| Autori principali: | Ma, Jie, Hu, Min, Wang, Pinghui, Sun, Wangchun, Song, Lingyun, Pei, Hongbin, Liu, Jun, Du, Youtian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Robust Visual Question Answering: Datasets, Methods, and Future Challenges
di: Ma, Jie, et al.
Pubblicazione: (2023)
di: Ma, Jie, et al.
Pubblicazione: (2023)
An Empirical Study for Representations of Videos in Video Question Answering via MLLMs
di: Li, Zhi, et al.
Pubblicazione: (2025)
di: Li, Zhi, et al.
Pubblicazione: (2025)
ZeShot-VQA: Zero-Shot Visual Question Answering Framework with Answer Mapping for Natural Disaster Damage Assessment
di: Karimi, Ehsan, et al.
Pubblicazione: (2025)
di: Karimi, Ehsan, et al.
Pubblicazione: (2025)
An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
di: Castrillón-Santana, Modesto, et al.
Pubblicazione: (2025)
di: Castrillón-Santana, Modesto, et al.
Pubblicazione: (2025)
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
di: Shah, Nisarg A., et al.
Pubblicazione: (2025)
di: Shah, Nisarg A., et al.
Pubblicazione: (2025)
Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering
di: Fan, Lin, et al.
Pubblicazione: (2026)
di: Fan, Lin, et al.
Pubblicazione: (2026)
Medico 2025: Visual Question Answering for Gastrointestinal Imaging
di: Gautam, Sushant, et al.
Pubblicazione: (2025)
di: Gautam, Sushant, et al.
Pubblicazione: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
Mixture of Rationale: Multi-Modal Reasoning Mixture for Visual Question Answering
di: Li, Tao, et al.
Pubblicazione: (2024)
di: Li, Tao, et al.
Pubblicazione: (2024)
The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering
di: Pandey, Anupam, et al.
Pubblicazione: (2025)
di: Pandey, Anupam, et al.
Pubblicazione: (2025)
Debate on Graph: a Flexible and Reliable Reasoning Framework for Large Language Models
di: Ma, Jie, et al.
Pubblicazione: (2024)
di: Ma, Jie, et al.
Pubblicazione: (2024)
VeS: Teaching Pixels to Listen Without Supervision
di: Raj, Sajay
Pubblicazione: (2025)
di: Raj, Sajay
Pubblicazione: (2025)
NAAQA: A Neural Architecture for Acoustic Question Answering
di: Abdelnour, Jerome, et al.
Pubblicazione: (2021)
di: Abdelnour, Jerome, et al.
Pubblicazione: (2021)
Semi-supervised Latent Disentangled Diffusion Model for Textile Pattern Generation
di: Hu, Chenggong, et al.
Pubblicazione: (2026)
di: Hu, Chenggong, et al.
Pubblicazione: (2026)
AVControl: Efficient Framework for Training Audio-Visual Controls
di: Ben-Yosef, Matan, et al.
Pubblicazione: (2026)
di: Ben-Yosef, Matan, et al.
Pubblicazione: (2026)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
di: Yim, Wen-wai, et al.
Pubblicazione: (2025)
di: Yim, Wen-wai, et al.
Pubblicazione: (2025)
FocusedAD: Character-centric Movie Audio Description
di: Ye, Xiaojun, et al.
Pubblicazione: (2025)
di: Ye, Xiaojun, et al.
Pubblicazione: (2025)
Tri-VQA: Triangular Reasoning Medical Visual Question Answering for Multi-Attribute Analysis
di: Fan, Lin, et al.
Pubblicazione: (2024)
di: Fan, Lin, et al.
Pubblicazione: (2024)
Visual Graph Question Answering with ASP and LLMs for Language Parsing
di: Bauer, Jakob Johannes, et al.
Pubblicazione: (2025)
di: Bauer, Jakob Johannes, et al.
Pubblicazione: (2025)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
di: Li, Danyang, et al.
Pubblicazione: (2025)
di: Li, Danyang, et al.
Pubblicazione: (2025)
MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition
di: Hu, Yudong, et al.
Pubblicazione: (2025)
di: Hu, Yudong, et al.
Pubblicazione: (2025)
SoccerRef-Agents: Multi-Agent System for Automated Soccer Refereeing
di: Meng, Zi, et al.
Pubblicazione: (2026)
di: Meng, Zi, et al.
Pubblicazione: (2026)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
di: Kashyap, Pankhi, et al.
Pubblicazione: (2024)
di: Kashyap, Pankhi, et al.
Pubblicazione: (2024)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
di: Li, Huibin, et al.
Pubblicazione: (2025)
di: Li, Huibin, et al.
Pubblicazione: (2025)
RAE-NWM: Navigation World Model in Dense Visual Representation Space
di: Zhang, Mingkun, et al.
Pubblicazione: (2026)
di: Zhang, Mingkun, et al.
Pubblicazione: (2026)
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
di: Balasubramanian, Sriram, et al.
Pubblicazione: (2025)
di: Balasubramanian, Sriram, et al.
Pubblicazione: (2025)
SUN Team's Contribution to ABAW 2024 Competition: Audio-visual Valence-Arousal Estimation and Expression Recognition
di: Dresvyanskiy, Denis, et al.
Pubblicazione: (2024)
di: Dresvyanskiy, Denis, et al.
Pubblicazione: (2024)
KARMA-MV: A Benchmark for Causal Question Answering on Music Videos
di: Ghosh, Archishman, et al.
Pubblicazione: (2026)
di: Ghosh, Archishman, et al.
Pubblicazione: (2026)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
di: Bian, Zhipeng, et al.
Pubblicazione: (2026)
di: Bian, Zhipeng, et al.
Pubblicazione: (2026)
Evaluating Perspectival Biases in Cross-Modal Retrieval
di: Saengsukhiran, Teerapol, et al.
Pubblicazione: (2025)
di: Saengsukhiran, Teerapol, et al.
Pubblicazione: (2025)
CoMoCAVs: Cohesive Decision-Guided Motion Planning for Connected and Autonomous Vehicles with Multi-Policy Reinforcement Learning
di: Hu, Pan
Pubblicazione: (2025)
di: Hu, Pan
Pubblicazione: (2025)
UAV-assisted Visual SLAM Generating Reconstructed 3D Scene Graphs in GPS-denied Environments
di: Radwan, Ahmed, et al.
Pubblicazione: (2024)
di: Radwan, Ahmed, et al.
Pubblicazione: (2024)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
di: Dai, Song, et al.
Pubblicazione: (2025)
di: Dai, Song, et al.
Pubblicazione: (2025)
A Vision-Language Model for Focal Liver Lesion Classification
di: Jian, Song, et al.
Pubblicazione: (2025)
di: Jian, Song, et al.
Pubblicazione: (2025)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
di: Yang, Shan
Pubblicazione: (2026)
di: Yang, Shan
Pubblicazione: (2026)
Biased Heritage: How Datasets Shape Models in Facial Expression Recognition
di: Dominguez-Catena, Iris, et al.
Pubblicazione: (2025)
di: Dominguez-Catena, Iris, et al.
Pubblicazione: (2025)
FlowDet: Overcoming Perspective and Scale Challenges in Real-Time End-to-End Traffic Detection
di: Wang, Zixing, et al.
Pubblicazione: (2025)
di: Wang, Zixing, et al.
Pubblicazione: (2025)
SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory
di: Alam, Samiul, et al.
Pubblicazione: (2026)
di: Alam, Samiul, et al.
Pubblicazione: (2026)
Learning Association via Track-Detection Matching for Multi-Object Tracking
di: Adžemović, Momir
Pubblicazione: (2025)
di: Adžemović, Momir
Pubblicazione: (2025)
DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
di: Deng, Pei, et al.
Pubblicazione: (2025)
di: Deng, Pei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Robust Visual Question Answering: Datasets, Methods, and Future Challenges
di: Ma, Jie, et al.
Pubblicazione: (2023) -
An Empirical Study for Representations of Videos in Video Question Answering via MLLMs
di: Li, Zhi, et al.
Pubblicazione: (2025) -
ZeShot-VQA: Zero-Shot Visual Question Answering Framework with Answer Mapping for Natural Disaster Damage Assessment
di: Karimi, Ehsan, et al.
Pubblicazione: (2025) -
An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
di: Castrillón-Santana, Modesto, et al.
Pubblicazione: (2025) -
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
di: Shah, Nisarg A., et al.
Pubblicazione: (2025)