CommVQA: Situating Visual Question Answering in Communicative Contexts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Naik, Nandita Shankar, Potts, Christopher, Kreiss, Elisa |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
von: Verma, Arnav, et al.
Veröffentlicht: (2025)
von: Verma, Arnav, et al.
Veröffentlicht: (2025)
VQA-MHUG: A Gaze Dataset to Study Multimodal Neural Attention in Visual Question Answering
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
Adversarial Attacks on VQA-NLE: Exposing and Alleviating Inconsistencies in Visual Question Answering Explanations
von: Yeh, Yahsin, et al.
Veröffentlicht: (2025)
von: Yeh, Yahsin, et al.
Veröffentlicht: (2025)
UCSF-PDGM-VQA: Visual Question Answering dataset for brain tumor MRI interpretation
von: Ghosh, Shiv, et al.
Veröffentlicht: (2026)
von: Ghosh, Shiv, et al.
Veröffentlicht: (2026)
Updating CLIP to Prefer Descriptions Over Captions
von: Zur, Amir, et al.
Veröffentlicht: (2024)
von: Zur, Amir, et al.
Veröffentlicht: (2024)
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts
von: Singh, Shubhankar, et al.
Veröffentlicht: (2024)
von: Singh, Shubhankar, et al.
Veröffentlicht: (2024)
Visual Robustness Benchmark for Visual Question Answering (VQA)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
Selectively Answering Visual Questions
von: Eisenschlos, Julian Martin, et al.
Veröffentlicht: (2024)
von: Eisenschlos, Julian Martin, et al.
Veröffentlicht: (2024)
RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
von: Butsanets, Léo, et al.
Veröffentlicht: (2025)
von: Butsanets, Léo, et al.
Veröffentlicht: (2025)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
BERT-VQA: Visual Question Answering on Plots
von: Vu, Tai, et al.
Veröffentlicht: (2025)
von: Vu, Tai, et al.
Veröffentlicht: (2025)
FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions
von: Inadumi, Shun, et al.
Veröffentlicht: (2024)
von: Inadumi, Shun, et al.
Veröffentlicht: (2024)
Advancing Vietnamese Visual Question Answering with Transformer and Convolutional Integration
von: Nguyen, Ngoc Son, et al.
Veröffentlicht: (2024)
von: Nguyen, Ngoc Son, et al.
Veröffentlicht: (2024)
Multimodal Integration of Human-Like Attention in Visual Question Answering
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
StackOverflowVQA: Stack Overflow Visual Question Answering Dataset
von: Mirzaei, Motahhare, et al.
Veröffentlicht: (2024)
von: Mirzaei, Motahhare, et al.
Veröffentlicht: (2024)
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering
von: Ma, Ziyu, et al.
Veröffentlicht: (2024)
von: Ma, Ziyu, et al.
Veröffentlicht: (2024)
Large Vision-Language Models for Remote Sensing Visual Question Answering
von: Siripong, Surasakdi, et al.
Veröffentlicht: (2024)
von: Siripong, Surasakdi, et al.
Veröffentlicht: (2024)
Design as Desired: Utilizing Visual Question Answering for Multimodal Pre-training
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
Computed Tomography Visual Question Answering with Cross-modal Feature Graphing
von: Tian, Yuanhe, et al.
Veröffentlicht: (2025)
von: Tian, Yuanhe, et al.
Veröffentlicht: (2025)
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
von: Awal, Rabiul, et al.
Veröffentlicht: (2023)
von: Awal, Rabiul, et al.
Veröffentlicht: (2023)
AIM: Asymmetric Information Masking for Visual Question Answering Continual Learning
von: Zhang, Peifeng, et al.
Veröffentlicht: (2026)
von: Zhang, Peifeng, et al.
Veröffentlicht: (2026)
Overcoming Language Priors for Visual Question Answering Based on Knowledge Distillation
von: Peng, Daowan, et al.
Veröffentlicht: (2025)
von: Peng, Daowan, et al.
Veröffentlicht: (2025)
KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
von: Hwang, Taebaek, et al.
Veröffentlicht: (2025)
von: Hwang, Taebaek, et al.
Veröffentlicht: (2025)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
von: Shrestha, Robik, et al.
Veröffentlicht: (2020)
von: Shrestha, Robik, et al.
Veröffentlicht: (2020)
DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes
von: Al-Mohannadi, Aisha, et al.
Veröffentlicht: (2026)
von: Al-Mohannadi, Aisha, et al.
Veröffentlicht: (2026)
MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
von: Nguyen, Hai-Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Hai-Dang, et al.
Veröffentlicht: (2025)
Towards Perceiving Small Visual Details in Zero-shot Visual Question Answering with Multimodal LLMs
von: Zhang, Jiarui, et al.
Veröffentlicht: (2023)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2023)
LaPA: Latent Prompt Assist Model For Medical Visual Question Answering
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
Fusion of Domain-Adapted Vision and Language Models for Medical Visual Question Answering
von: Ha, Cuong Nhat, et al.
Veröffentlicht: (2024)
von: Ha, Cuong Nhat, et al.
Veröffentlicht: (2024)
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
von: Kuang, Jiayi, et al.
Veröffentlicht: (2024)
von: Kuang, Jiayi, et al.
Veröffentlicht: (2024)
ChitroJera: A Regionally Relevant Visual Question Answering Dataset for Bangla
von: Barua, Deeparghya Dutta, et al.
Veröffentlicht: (2024)
von: Barua, Deeparghya Dutta, et al.
Veröffentlicht: (2024)
How Good LLMs Are at Answering Bangla Medical Visual Questions? Dataset and Benchmarking
von: Ahmed, Rafid, et al.
Veröffentlicht: (2026)
von: Ahmed, Rafid, et al.
Veröffentlicht: (2026)
Grounding Chest X-Ray Visual Question Answering with Generated Radiology Reports
von: Serra, Francesco Dalla, et al.
Veröffentlicht: (2025)
von: Serra, Francesco Dalla, et al.
Veröffentlicht: (2025)
Hallucination Benchmark in Medical Visual Question Answering
von: Wu, Jinge, et al.
Veröffentlicht: (2024)
von: Wu, Jinge, et al.
Veröffentlicht: (2024)
Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQA
von: Mo, Wentao, et al.
Veröffentlicht: (2024)
von: Mo, Wentao, et al.
Veröffentlicht: (2024)
Multi-Agent VQA: Exploring Multi-Agent Foundation Models in Zero-Shot Visual Question Answering
von: Jiang, Bowen, et al.
Veröffentlicht: (2024)
von: Jiang, Bowen, et al.
Veröffentlicht: (2024)
VQA$^2$: Visual Question Answering for Video Quality Assessment
von: Jia, Ziheng, et al.
Veröffentlicht: (2024)
von: Jia, Ziheng, et al.
Veröffentlicht: (2024)
Integrating Image Features with Convolutional Sequence-to-sequence Network for Multilingual Visual Question Answering
von: Thai, Triet Minh, et al.
Veröffentlicht: (2023)
von: Thai, Triet Minh, et al.
Veröffentlicht: (2023)
MOTOR: Multimodal Optimal Transport via Grounded Retrieval in Medical Visual Question Answering
von: Shaaban, Mai A., et al.
Veröffentlicht: (2025)
von: Shaaban, Mai A., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models
von: Verma, Arnav, et al.
Veröffentlicht: (2025) -
VQA-MHUG: A Gaze Dataset to Study Multimodal Neural Attention in Visual Question Answering
von: Sood, Ekta, et al.
Veröffentlicht: (2021) -
Adversarial Attacks on VQA-NLE: Exposing and Alleviating Inconsistencies in Visual Question Answering Explanations
von: Yeh, Yahsin, et al.
Veröffentlicht: (2025) -
UCSF-PDGM-VQA: Visual Question Answering dataset for brain tumor MRI interpretation
von: Ghosh, Shiv, et al.
Veröffentlicht: (2026) -
Updating CLIP to Prefer Descriptions Over Captions
von: Zur, Amir, et al.
Veröffentlicht: (2024)