Enhancing Visual Question Answering through Question-Driven Image Captions as Prompts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Özdemir, Övgü, Akagündüz, Erdem |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering
von: Anaissi, Ali, et al.
Veröffentlicht: (2025)
von: Anaissi, Ali, et al.
Veröffentlicht: (2025)
Knowledge Detection by Relevant Question and Image Attributes in Visual Question Answering
von: Ahir, Param, et al.
Veröffentlicht: (2023)
von: Ahir, Param, et al.
Veröffentlicht: (2023)
SCRA-VQA: Summarized Caption-Rerank for Augmented Large Language Models in Visual Question Answering
von: Zhang, Yan, et al.
Veröffentlicht: (2025)
von: Zhang, Yan, et al.
Veröffentlicht: (2025)
RS-MoE: A Vision-Language Model with Mixture of Experts for Remote Sensing Image Captioning and Visual Question Answering
von: Lin, Hui, et al.
Veröffentlicht: (2024)
von: Lin, Hui, et al.
Veröffentlicht: (2024)
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering
von: Ukai, Mahiro, et al.
Veröffentlicht: (2024)
von: Ukai, Mahiro, et al.
Veröffentlicht: (2024)
TM-PATHVQA:90000+ Textless Multilingual Questions for Medical Visual Question Answering
von: Rajkhowa, Tonmoy, et al.
Veröffentlicht: (2024)
von: Rajkhowa, Tonmoy, et al.
Veröffentlicht: (2024)
VoQA: Visual-only Question Answering
von: An, Jianing, et al.
Veröffentlicht: (2025)
von: An, Jianing, et al.
Veröffentlicht: (2025)
Object Attribute Matters in Visual Question Answering
von: Li, Peize, et al.
Veröffentlicht: (2023)
von: Li, Peize, et al.
Veröffentlicht: (2023)
MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering
von: Li, Xu, et al.
Veröffentlicht: (2025)
von: Li, Xu, et al.
Veröffentlicht: (2025)
Near-Infrared and Low-Rank Adaptation of Vision Transformers in Remote Sensing
von: Ulku, Irem, et al.
Veröffentlicht: (2024)
von: Ulku, Irem, et al.
Veröffentlicht: (2024)
Robust Multispectral Semantic Segmentation under Missing or Full Modalities via Structured Latent Projection
von: Ulku, Irem, et al.
Veröffentlicht: (2026)
von: Ulku, Irem, et al.
Veröffentlicht: (2026)
Latent Space Guided Scenario Sampling for Multimodal Segmentation Under Missing Modalities
von: Ulku, Irem, et al.
Veröffentlicht: (2026)
von: Ulku, Irem, et al.
Veröffentlicht: (2026)
Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question Answering
von: Marouf, Imad Eddine, et al.
Veröffentlicht: (2025)
von: Marouf, Imad Eddine, et al.
Veröffentlicht: (2025)
CluMo: Cluster-based Modality Fusion Prompt for Continual Learning in Visual Question Answering
von: Cai, Yuliang, et al.
Veröffentlicht: (2024)
von: Cai, Yuliang, et al.
Veröffentlicht: (2024)
Cause-Effect Driven Optimization for Robust Medical Visual Question Answering with Language Biases
von: Zhu, Huanjia, et al.
Veröffentlicht: (2025)
von: Zhu, Huanjia, et al.
Veröffentlicht: (2025)
WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
von: Chen, Pingyi, et al.
Veröffentlicht: (2024)
von: Chen, Pingyi, et al.
Veröffentlicht: (2024)
Free Form Medical Visual Question Answering in Radiology
von: Narayanan, Abhishek, et al.
Veröffentlicht: (2024)
von: Narayanan, Abhishek, et al.
Veröffentlicht: (2024)
Saliency Guided Longitudinal Medical Visual Question Answering
von: Wu, Jialin, et al.
Veröffentlicht: (2025)
von: Wu, Jialin, et al.
Veröffentlicht: (2025)
Multi-Sourced Compositional Generalization in Visual Question Answering
von: Li, Chuanhao, et al.
Veröffentlicht: (2025)
von: Li, Chuanhao, et al.
Veröffentlicht: (2025)
FuseFormer: A Transformer for Visual and Thermal Image Fusion
von: Erdogan, Aytekin, et al.
Veröffentlicht: (2024)
von: Erdogan, Aytekin, et al.
Veröffentlicht: (2024)
Hallucination Benchmark in Medical Visual Question Answering
von: Wu, Jinge, et al.
Veröffentlicht: (2024)
von: Wu, Jinge, et al.
Veröffentlicht: (2024)
DermaBench: A Clinician-Annotated Benchmark Dataset for Dermatology Visual Question Answering and Reasoning
von: Yilmaz, Abdurrahim, et al.
Veröffentlicht: (2026)
von: Yilmaz, Abdurrahim, et al.
Veröffentlicht: (2026)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
von: Oh, Ju-Young, et al.
Veröffentlicht: (2025)
von: Oh, Ju-Young, et al.
Veröffentlicht: (2025)
FaithSCAN: Model-Driven Single-Pass Hallucination Detection for Faithful Visual Question Answering
von: Tong, Chaodong, et al.
Veröffentlicht: (2026)
von: Tong, Chaodong, et al.
Veröffentlicht: (2026)
Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering
von: Lim, Youngsun, et al.
Veröffentlicht: (2024)
von: Lim, Youngsun, et al.
Veröffentlicht: (2024)
VQA$^2$: Visual Question Answering for Video Quality Assessment
von: Jia, Ziheng, et al.
Veröffentlicht: (2024)
von: Jia, Ziheng, et al.
Veröffentlicht: (2024)
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
von: Wieczorek, Tobias Jan, et al.
Veröffentlicht: (2025)
von: Wieczorek, Tobias Jan, et al.
Veröffentlicht: (2025)
Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering
von: Zhang, Zhengxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengxuan, et al.
Veröffentlicht: (2025)
Wasserstein Equilibrium Decoding for Reliable Medical Visual Question Answering
von: Hagen, Luca, et al.
Veröffentlicht: (2026)
von: Hagen, Luca, et al.
Veröffentlicht: (2026)
Location-Aware Pretraining for Medical Difference Visual Question Answering
von: Musinguzi, Denis, et al.
Veröffentlicht: (2026)
von: Musinguzi, Denis, et al.
Veröffentlicht: (2026)
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
von: Gupta, Akash, et al.
Veröffentlicht: (2025)
von: Gupta, Akash, et al.
Veröffentlicht: (2025)
MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
von: Xi, Suyang, et al.
Veröffentlicht: (2026)
von: Xi, Suyang, et al.
Veröffentlicht: (2026)
Parameter-Efficient VLMs for Gastrointestinal Endoscopy: Medical Image Generation and Clinical Visual Question Answering
von: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Veröffentlicht: (2026)
von: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Veröffentlicht: (2026)
LingoQA: Visual Question Answering for Autonomous Driving
von: Marcu, Ana-Maria, et al.
Veröffentlicht: (2023)
von: Marcu, Ana-Maria, et al.
Veröffentlicht: (2023)
Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question Answering
von: Shen, Ruoyue, et al.
Veröffentlicht: (2024)
von: Shen, Ruoyue, et al.
Veröffentlicht: (2024)
IIU: Independent Inference Units for Knowledge-based Visual Question Answering
von: Li, Yili, et al.
Veröffentlicht: (2024)
von: Li, Yili, et al.
Veröffentlicht: (2024)
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
von: Choi, Changin, et al.
Veröffentlicht: (2025)
von: Choi, Changin, et al.
Veröffentlicht: (2025)
Hierarchical Modeling for Medical Visual Question Answering with Cross-Attention Fusion
von: Zhang, Junkai, et al.
Veröffentlicht: (2025)
von: Zhang, Junkai, et al.
Veröffentlicht: (2025)
Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring
von: Vu, Sinh Trong, et al.
Veröffentlicht: (2025)
von: Vu, Sinh Trong, et al.
Veröffentlicht: (2025)
Thought-For-Food: Reasoning Chain Induced Food Visual Question Answering
von: Jain, Riddhi, et al.
Veröffentlicht: (2025)
von: Jain, Riddhi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering
von: Anaissi, Ali, et al.
Veröffentlicht: (2025) -
Knowledge Detection by Relevant Question and Image Attributes in Visual Question Answering
von: Ahir, Param, et al.
Veröffentlicht: (2023) -
SCRA-VQA: Summarized Caption-Rerank for Augmented Large Language Models in Visual Question Answering
von: Zhang, Yan, et al.
Veröffentlicht: (2025) -
RS-MoE: A Vision-Language Model with Mixture of Experts for Remote Sensing Image Captioning and Visual Question Answering
von: Lin, Hui, et al.
Veröffentlicht: (2024) -
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering
von: Ukai, Mahiro, et al.
Veröffentlicht: (2024)