An Evaluation of GPT-4V and Gemini in Online VQA
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Mengchen, Chen, Chongyan, Gurari, Danna |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fully Authentic Visual Question Answering Dataset from Online Communities
by: Chen, Chongyan, et al.
Published: (2023)
by: Chen, Chongyan, et al.
Published: (2023)
Acknowledging Focus Ambiguity in Visual Questions
by: Chen, Chongyan, et al.
Published: (2025)
by: Chen, Chongyan, et al.
Published: (2025)
Logit-Based Losses Limit the Effectiveness of Feature Knowledge Distillation
by: Cooper, Nicholas, et al.
Published: (2025)
by: Cooper, Nicholas, et al.
Published: (2025)
Visual Reasoning Evaluation of Grok, Deepseek Janus, Gemini, Qwen, Mistral, and ChatGPT
by: Jegham, Nidhal, et al.
Published: (2025)
by: Jegham, Nidhal, et al.
Published: (2025)
Is ChatGPT-5 Ready for Mammogram VQA?
by: Li, Qiang, et al.
Published: (2025)
by: Li, Qiang, et al.
Published: (2025)
GPT as Psychologist? Preliminary Evaluations for GPT-4V on Visual Affective Computing
by: Lu, Hao, et al.
Published: (2024)
by: Lu, Hao, et al.
Published: (2024)
Collecting Consistently High Quality Object Tracks with Minimal Human Involvement by Using Self-Supervised Learning to Detect Tracker Errors
by: Anjum, Samreen, et al.
Published: (2024)
by: Anjum, Samreen, et al.
Published: (2024)
A Systematic Evaluation of GPT-4V's Multimodal Capability for Medical Image Analysis
by: Li, Yingshu, et al.
Published: (2023)
by: Li, Yingshu, et al.
Published: (2023)
PartStickers: Generating Parts of Objects for Rapid Prototyping
by: Zhou, Mo, et al.
Published: (2025)
by: Zhou, Mo, et al.
Published: (2025)
Can GPT-4o mini and Gemini 2.0 Flash Predict Fine-Grained Fashion Product Attributes? A Zero-Shot Analysis
by: Shukla, Shubham, et al.
Published: (2025)
by: Shukla, Shubham, et al.
Published: (2025)
Long-Form Answers to Visual Questions from Blind and Low Vision People
by: Huh, Mina, et al.
Published: (2024)
by: Huh, Mina, et al.
Published: (2024)
Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs
by: Kurz, Paul Jonas, et al.
Published: (2026)
by: Kurz, Paul Jonas, et al.
Published: (2026)
Assessing Greenspace Attractiveness with ChatGPT, Claude, and Gemini: Do AI Models Reflect Human Perceptions?
by: Malekzadeh, Milad, et al.
Published: (2025)
by: Malekzadeh, Milad, et al.
Published: (2025)
R-LLaVA: Improving Med-VQA Understanding through Visual Region of Interest
by: Chen, Xupeng, et al.
Published: (2024)
by: Chen, Xupeng, et al.
Published: (2024)
CollagePrompt: A Benchmark for Budget-Friendly Visual Recognition with GPT-4V
by: Xu, Siyu, et al.
Published: (2024)
by: Xu, Siyu, et al.
Published: (2024)
MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Space
by: Singh, Anshul, et al.
Published: (2025)
by: Singh, Anshul, et al.
Published: (2025)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
by: Zhang, Mengchen, et al.
Published: (2025)
by: Zhang, Mengchen, et al.
Published: (2025)
InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts
by: Xie, Tianchi, et al.
Published: (2025)
by: Xie, Tianchi, et al.
Published: (2025)
MISS: A Generative Pretraining and Finetuning Approach for Med-VQA
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
Advancing Surgical VQA with Scene Graph Knowledge
by: Yuan, Kun, et al.
Published: (2023)
by: Yuan, Kun, et al.
Published: (2023)
VQA$^2$: Visual Question Answering for Video Quality Assessment
by: Jia, Ziheng, et al.
Published: (2024)
by: Jia, Ziheng, et al.
Published: (2024)
Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence
by: Wang, Xingyue, et al.
Published: (2026)
by: Wang, Xingyue, et al.
Published: (2026)
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
by: Li, Yanwei, et al.
Published: (2024)
by: Li, Yanwei, et al.
Published: (2024)
WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
by: Chen, Pingyi, et al.
Published: (2024)
by: Chen, Pingyi, et al.
Published: (2024)
AlignGemini: Generalizable AI-Generated Image Detection Through Task-Model Alignment
by: Chen, Ruoxin, et al.
Published: (2025)
by: Chen, Ruoxin, et al.
Published: (2025)
GEMeX-RMCoT: An Enhanced Med-VQA Dataset for Region-Aware Multimodal Chain-of-Thought Reasoning
by: Liu, Bo, et al.
Published: (2025)
by: Liu, Bo, et al.
Published: (2025)
KNVQA: A Benchmark for evaluation knowledge-based VQA
by: Cheng, Sirui, et al.
Published: (2023)
by: Cheng, Sirui, et al.
Published: (2023)
Is it safe to cross? Interpretable Risk Assessment with GPT-4V for Safety-Aware Street Crossing
by: Hwang, Hochul, et al.
Published: (2024)
by: Hwang, Hochul, et al.
Published: (2024)
VQA-Levels: A Hierarchical Approach for Classifying Questions in VQA
by: Madaka, Madhuri Latha, et al.
Published: (2025)
by: Madaka, Madhuri Latha, et al.
Published: (2025)
A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5
by: Ma, Xingjun, et al.
Published: (2026)
by: Ma, Xingjun, et al.
Published: (2026)
VSA4VQA: Scaling a Vector Symbolic Architecture to Visual Question Answering on Natural Images
by: Penzkofer, Anna, et al.
Published: (2024)
by: Penzkofer, Anna, et al.
Published: (2024)
DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents
by: Chen, Yixiong, et al.
Published: (2026)
by: Chen, Yixiong, et al.
Published: (2026)
GPT-4V-AD: Exploring Grounding Potential of VQA-oriented GPT-4V for Zero-shot Anomaly Detection
by: Zhang, Jiangning, et al.
Published: (2023)
by: Zhang, Jiangning, et al.
Published: (2023)
Evaluating Gemini Robotics Policies in a Veo World Simulator
by: Gemini Robotics Team, et al.
Published: (2025)
by: Gemini Robotics Team, et al.
Published: (2025)
ARIAL: An Agentic Framework for Document VQA with Precise Answer Localization
by: Mohammadshirazi, Ahmad, et al.
Published: (2025)
by: Mohammadshirazi, Ahmad, et al.
Published: (2025)
Box-QAymo: Box-Referring VQA Dataset for Autonomous Driving
by: Etchegaray, Djamahl, et al.
Published: (2025)
by: Etchegaray, Djamahl, et al.
Published: (2025)
R^3-VQA: "Read the Room" by Video Social Reasoning
by: Niu, Lixing, et al.
Published: (2025)
by: Niu, Lixing, et al.
Published: (2025)
Better Supervised Fine-tuning for VQA: Integer-Only Loss
by: Qian, Baihong, et al.
Published: (2025)
by: Qian, Baihong, et al.
Published: (2025)
Gemini: A Family of Highly Capable Multimodal Models
by: Gemini Team, et al.
Published: (2023)
by: Gemini Team, et al.
Published: (2023)
Evaluation of GPT-4o and GPT-4o-mini's Vision Capabilities for Compositional Analysis from Dried Solution Drops
by: Dangi, Deven B., et al.
Published: (2024)
by: Dangi, Deven B., et al.
Published: (2024)
Similar Items
-
Fully Authentic Visual Question Answering Dataset from Online Communities
by: Chen, Chongyan, et al.
Published: (2023) -
Acknowledging Focus Ambiguity in Visual Questions
by: Chen, Chongyan, et al.
Published: (2025) -
Logit-Based Losses Limit the Effectiveness of Feature Knowledge Distillation
by: Cooper, Nicholas, et al.
Published: (2025) -
Visual Reasoning Evaluation of Grok, Deepseek Janus, Gemini, Qwen, Mistral, and ChatGPT
by: Jegham, Nidhal, et al.
Published: (2025) -
Is ChatGPT-5 Ready for Mammogram VQA?
by: Li, Qiang, et al.
Published: (2025)