From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities
Fuente:
arXiv
Saved in:
| Main Authors: | Ishmam, Md Farhan, Shovon, Md Sakib Hossain, Mridha, M. F., Dey, Nilanjan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Robustness Benchmark for Visual Question Answering (VQA)
by: Ishmam, Md Farhan, et al.
Published: (2024)
by: Ishmam, Md Farhan, et al.
Published: (2024)
ChitroJera: A Regionally Relevant Visual Question Answering Dataset for Bangla
by: Barua, Deeparghya Dutta, et al.
Published: (2024)
by: Barua, Deeparghya Dutta, et al.
Published: (2024)
MobilePlantViT: A Mobile-friendly Hybrid ViT for Generalized Plant Disease Image Classification
by: Tonmoy, Moshiur Rahman, et al.
Published: (2025)
by: Tonmoy, Moshiur Rahman, et al.
Published: (2025)
PlantVillageVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science
by: Sakib, Syed Nazmus, et al.
Published: (2025)
by: Sakib, Syed Nazmus, et al.
Published: (2025)
Soybean Disease Detection via Interpretable Hybrid CNN-GNN: Integrating MobileNetV2 and GraphSAGE with Cross-Modal Attention
by: Jahin, Md Abrar, et al.
Published: (2025)
by: Jahin, Md Abrar, et al.
Published: (2025)
KDC-Diff: A Latent-Aware Diffusion Model with Knowledge Retention for Memory-Efficient Image Generation
by: Borno, Md. Naimur Asif, et al.
Published: (2025)
by: Borno, Md. Naimur Asif, et al.
Published: (2025)
A Two-Stage Multitask Vision-Language Framework for Explainable Crop Disease Visual Question Answering
by: Hossain, Md. Zahid, et al.
Published: (2026)
by: Hossain, Md. Zahid, et al.
Published: (2026)
TimeWarp: Evaluating Web Agents by Revisiting the Past
by: Ishmam, Md Farhan, et al.
Published: (2026)
by: Ishmam, Md Farhan, et al.
Published: (2026)
How Good LLMs Are at Answering Bangla Medical Visual Questions? Dataset and Benchmarking
by: Ahmed, Rafid, et al.
Published: (2026)
by: Ahmed, Rafid, et al.
Published: (2026)
Co-AttenDWG: Co-Attentive Dimension-Wise Gating and Expert Fusion for Multi-Modal Offensive Content Detection
by: Hossain, Md. Mithun, et al.
Published: (2025)
by: Hossain, Md. Mithun, et al.
Published: (2025)
Prompting with Sign Parameters for Low-resource Sign Language Instruction Generation
by: Tariquzzaman, Md, et al.
Published: (2025)
by: Tariquzzaman, Md, et al.
Published: (2025)
BERT-VQA: Visual Question Answering on Plots
by: Vu, Tai, et al.
Published: (2025)
by: Vu, Tai, et al.
Published: (2025)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
by: Zhang, Xiaoman, et al.
Published: (2023)
by: Zhang, Xiaoman, et al.
Published: (2023)
Semantic Shield: Defending Vision-Language Models Against Backdooring and Poisoning via Fine-grained Knowledge Alignment
by: Ishmam, Alvi Md, et al.
Published: (2024)
by: Ishmam, Alvi Md, et al.
Published: (2024)
StackOverflowVQA: Stack Overflow Visual Question Answering Dataset
by: Mirzaei, Motahhare, et al.
Published: (2024)
by: Mirzaei, Motahhare, et al.
Published: (2024)
WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
by: Chen, Pingyi, et al.
Published: (2024)
by: Chen, Pingyi, et al.
Published: (2024)
PitVQA: Image-grounded Text Embedding LLM for Visual Question Answering in Pituitary Surgery
by: He, Runlong, et al.
Published: (2024)
by: He, Runlong, et al.
Published: (2024)
Personalized Federated Segmentation with Shared Feature Aggregation and Boundary-Focused Calibration
by: Tashdeed, Ishmam, et al.
Published: (2025)
by: Tashdeed, Ishmam, et al.
Published: (2025)
CommVQA: Situating Visual Question Answering in Communicative Contexts
by: Naik, Nandita Shankar, et al.
Published: (2024)
by: Naik, Nandita Shankar, et al.
Published: (2024)
VQA$^2$: Visual Question Answering for Video Quality Assessment
by: Jia, Ziheng, et al.
Published: (2024)
by: Jia, Ziheng, et al.
Published: (2024)
DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes
by: Al-Mohannadi, Aisha, et al.
Published: (2026)
by: Al-Mohannadi, Aisha, et al.
Published: (2026)
MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
by: Nguyen, Hai-Dang, et al.
Published: (2025)
by: Nguyen, Hai-Dang, et al.
Published: (2025)
Transformation of Biological Networks into Images via Semantic Cartography for Visual Interpretation and Scalable Deep Analysis
by: Mostafa, Sakib, et al.
Published: (2025)
by: Mostafa, Sakib, et al.
Published: (2025)
Reliable Deep Learning for Small-Scale Classifications: Experiments on Real-World Image Datasets from Bangladesh
by: Suny, Alfe, et al.
Published: (2026)
by: Suny, Alfe, et al.
Published: (2026)
Analyzing the Sensitivity of Vision Language Models in Visual Question Answering
by: Shah, Monika, et al.
Published: (2025)
by: Shah, Monika, et al.
Published: (2025)
Decentralized LoRA augmented transformer with multi-scale feature learning for secured eye diagnosis
by: Borno, Md. Naimur Asif, et al.
Published: (2025)
by: Borno, Md. Naimur Asif, et al.
Published: (2025)
Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring
by: Vu, Sinh Trong, et al.
Published: (2025)
by: Vu, Sinh Trong, et al.
Published: (2025)
A Comprehensive Survey on Visual Question Answering Datasets and Algorithms
by: Kabir, Raihan, et al.
Published: (2024)
by: Kabir, Raihan, et al.
Published: (2024)
SCRA-VQA: Summarized Caption-Rerank for Augmented Large Language Models in Visual Question Answering
by: Zhang, Yan, et al.
Published: (2025)
by: Zhang, Yan, et al.
Published: (2025)
ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering
by: Tran, Duong T., et al.
Published: (2025)
by: Tran, Duong T., et al.
Published: (2025)
MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question Answering
by: Li, Zhifei, et al.
Published: (2026)
by: Li, Zhifei, et al.
Published: (2026)
Detect2Interact: Localizing Object Key Field in Visual Question Answering (VQA) with LLMs
by: Wang, Jialou, et al.
Published: (2024)
by: Wang, Jialou, et al.
Published: (2024)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
by: Zhang, Chengyi, et al.
Published: (2026)
by: Zhang, Chengyi, et al.
Published: (2026)
CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
by: Han, Hongyong, et al.
Published: (2025)
by: Han, Hongyong, et al.
Published: (2025)
VSA4VQA: Scaling a Vector Symbolic Architecture to Visual Question Answering on Natural Images
by: Penzkofer, Anna, et al.
Published: (2024)
by: Penzkofer, Anna, et al.
Published: (2024)
Diffusion-Based Approaches in Medical Image Generation and Analysis
by: Nafi, Abdullah al Nomaan, et al.
Published: (2024)
by: Nafi, Abdullah al Nomaan, et al.
Published: (2024)
Single-Channel Tissue Segmentation via Cross-Modal Distillation from Foundation Models
by: Mohammad, Sakib, et al.
Published: (2026)
by: Mohammad, Sakib, et al.
Published: (2026)
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
by: Guan, Runwei, et al.
Published: (2025)
by: Guan, Runwei, et al.
Published: (2025)
AutoVQA-G: Self-Improving Agentic Framework for Automated Visual Question Answering and Grounding Annotation
by: Hu, Rongsheng, et al.
Published: (2026)
by: Hu, Rongsheng, et al.
Published: (2026)
ConMatFormer: A Multi-attention and Transformer Integrated ConvNext based Deep Learning Model for Enhanced Diabetic Foot Ulcer Classification
by: Rifat, Raihan Ahamed, et al.
Published: (2025)
by: Rifat, Raihan Ahamed, et al.
Published: (2025)
Similar Items
-
Visual Robustness Benchmark for Visual Question Answering (VQA)
by: Ishmam, Md Farhan, et al.
Published: (2024) -
ChitroJera: A Regionally Relevant Visual Question Answering Dataset for Bangla
by: Barua, Deeparghya Dutta, et al.
Published: (2024) -
MobilePlantViT: A Mobile-friendly Hybrid ViT for Generalized Plant Disease Image Classification
by: Tonmoy, Moshiur Rahman, et al.
Published: (2025) -
PlantVillageVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science
by: Sakib, Syed Nazmus, et al.
Published: (2025) -
Soybean Disease Detection via Interpretable Hybrid CNN-GNN: Integrating MobileNetV2 and GraphSAGE with Cross-Modal Attention
by: Jahin, Md Abrar, et al.
Published: (2025)