Evaluating Variance in Visual Question Answering Benchmarks
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | SR, Nikitha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DoPTA: Improving Document Layout Analysis using Patch-Text Alignment
von: SR, Nikitha, et al.
Veröffentlicht: (2024)
von: SR, Nikitha, et al.
Veröffentlicht: (2024)
Visual Robustness Benchmark for Visual Question Answering (VQA)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
von: Tang, Jingqun, et al.
Veröffentlicht: (2024)
von: Tang, Jingqun, et al.
Veröffentlicht: (2024)
Hallucination Benchmark in Medical Visual Question Answering
von: Wu, Jinge, et al.
Veröffentlicht: (2024)
von: Wu, Jinge, et al.
Veröffentlicht: (2024)
Towards Flexible Evaluation for Generative Visual Question Answering
von: Ji, Huishan, et al.
Veröffentlicht: (2024)
von: Ji, Huishan, et al.
Veröffentlicht: (2024)
DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes
von: Al-Mohannadi, Aisha, et al.
Veröffentlicht: (2026)
von: Al-Mohannadi, Aisha, et al.
Veröffentlicht: (2026)
Selectively Answering Visual Questions
von: Eisenschlos, Julian Martin, et al.
Veröffentlicht: (2024)
von: Eisenschlos, Julian Martin, et al.
Veröffentlicht: (2024)
MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
von: Song, Seokwon, et al.
Veröffentlicht: (2025)
von: Song, Seokwon, et al.
Veröffentlicht: (2025)
Dataset and Benchmark for Urdu Natural Scenes Text Detection, Recognition and Visual Question Answering
von: Maryam, Hiba, et al.
Veröffentlicht: (2024)
von: Maryam, Hiba, et al.
Veröffentlicht: (2024)
Visually Interpretable Subtask Reasoning for Visual Question Answering
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
von: Cheng, Yu, et al.
Veröffentlicht: (2025)
Targeted Visual Prompting for Medical Visual Question Answering
von: Tascon-Morales, Sergio, et al.
Veröffentlicht: (2024)
von: Tascon-Morales, Sergio, et al.
Veröffentlicht: (2024)
Question-Aware Gaussian Experts for Audio-Visual Question Answering
von: Kim, Hongyeob, et al.
Veröffentlicht: (2025)
von: Kim, Hongyeob, et al.
Veröffentlicht: (2025)
HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs
von: SR, Nikitha, et al.
Veröffentlicht: (2025)
von: SR, Nikitha, et al.
Veröffentlicht: (2025)
Questioning the Stability of Visual Question Answering
von: Rosenfeld, Amir, et al.
Veröffentlicht: (2025)
von: Rosenfeld, Amir, et al.
Veröffentlicht: (2025)
How Good LLMs Are at Answering Bangla Medical Visual Questions? Dataset and Benchmarking
von: Ahmed, Rafid, et al.
Veröffentlicht: (2026)
von: Ahmed, Rafid, et al.
Veröffentlicht: (2026)
Multimodal Rationales for Explainable Visual Question Answering
von: Li, Kun, et al.
Veröffentlicht: (2024)
von: Li, Kun, et al.
Veröffentlicht: (2024)
ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering
von: Tran, Duong T., et al.
Veröffentlicht: (2025)
von: Tran, Duong T., et al.
Veröffentlicht: (2025)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
von: Zhang, Chengyi, et al.
Veröffentlicht: (2026)
von: Zhang, Chengyi, et al.
Veröffentlicht: (2026)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
von: Romero, David, et al.
Veröffentlicht: (2024)
von: Romero, David, et al.
Veröffentlicht: (2024)
Evaluating Zero-Shot GPT-4V Performance on 3D Visual Question Answering Benchmarks
von: Singh, Simranjit, et al.
Veröffentlicht: (2024)
von: Singh, Simranjit, et al.
Veröffentlicht: (2024)
VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering
von: Wang, Yanling, et al.
Veröffentlicht: (2025)
von: Wang, Yanling, et al.
Veröffentlicht: (2025)
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario
von: Qian, Tianwen, et al.
Veröffentlicht: (2023)
von: Qian, Tianwen, et al.
Veröffentlicht: (2023)
DriveLM: Driving with Graph Visual Question Answering
von: Sima, Chonghao, et al.
Veröffentlicht: (2023)
von: Sima, Chonghao, et al.
Veröffentlicht: (2023)
Object Retrieval for Visual Question Answering with Outside Knowledge
von: Kan, Shichao, et al.
Veröffentlicht: (2024)
von: Kan, Shichao, et al.
Veröffentlicht: (2024)
Breaking the Visual Shortcuts in Multimodal Knowledge-Based Visual Question Answering
von: Lee, Dosung, et al.
Veröffentlicht: (2025)
von: Lee, Dosung, et al.
Veröffentlicht: (2025)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
von: Chintapatla, Ishant, et al.
Veröffentlicht: (2025)
von: Chintapatla, Ishant, et al.
Veröffentlicht: (2025)
SEA: Evaluating Sketch Abstraction Efficiency via Element-level Commonsense Visual Question Answering
von: Park, Jiho, et al.
Veröffentlicht: (2026)
von: Park, Jiho, et al.
Veröffentlicht: (2026)
Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark
von: Nguyen, Hieu Minh, et al.
Veröffentlicht: (2025)
von: Nguyen, Hieu Minh, et al.
Veröffentlicht: (2025)
ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images
von: Pham, Huy Quang, et al.
Veröffentlicht: (2024)
von: Pham, Huy Quang, et al.
Veröffentlicht: (2024)
QIRL: Boosting Visual Question Answering via Optimized Question-Image Relation Learning
von: Xu, Quanxing, et al.
Veröffentlicht: (2025)
von: Xu, Quanxing, et al.
Veröffentlicht: (2025)
VQ-VA World: Towards High-Quality Visual Question-Visual Answering
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
von: Gou, Chenhui, et al.
Veröffentlicht: (2025)
VIHD: Visual Intervention-based Hallucination Detection for Medical Visual Question Answering
von: Chen, Jiayi, et al.
Veröffentlicht: (2026)
von: Chen, Jiayi, et al.
Veröffentlicht: (2026)
Visual Question Answering on Multiple Remote Sensing Image Modalities
von: Boussaid, Hichem, et al.
Veröffentlicht: (2025)
von: Boussaid, Hichem, et al.
Veröffentlicht: (2025)
Reconstruction as a Bridge for Event-Based Visual Question Answering
von: Lou, Hanyue, et al.
Veröffentlicht: (2025)
von: Lou, Hanyue, et al.
Veröffentlicht: (2025)
Object-centric Video Question Answering with Visual Grounding and Referring
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
Analyzing the Sensitivity of Vision Language Models in Visual Question Answering
von: Shah, Monika, et al.
Veröffentlicht: (2025)
von: Shah, Monika, et al.
Veröffentlicht: (2025)
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
von: Xue, Junxiao, et al.
Veröffentlicht: (2024)
von: Xue, Junxiao, et al.
Veröffentlicht: (2024)
A Comprehensive Survey on Visual Question Answering Datasets and Algorithms
von: Kabir, Raihan, et al.
Veröffentlicht: (2024)
von: Kabir, Raihan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DoPTA: Improving Document Layout Analysis using Patch-Text Alignment
von: SR, Nikitha, et al.
Veröffentlicht: (2024) -
Visual Robustness Benchmark for Visual Question Answering (VQA)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024) -
MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
von: Tang, Jingqun, et al.
Veröffentlicht: (2024) -
Hallucination Benchmark in Medical Visual Question Answering
von: Wu, Jinge, et al.
Veröffentlicht: (2024) -
Towards Flexible Evaluation for Generative Visual Question Answering
von: Ji, Huishan, et al.
Veröffentlicht: (2024)