Attribute Diversity Determines the Systematicity Gap in VQA
Fuente:
arXiv
Saved in:
| Main Authors: | Berlot-Attwell, Ian, Agrawal, Kumar Krishna, Carrell, A. Michael, Sharma, Yash, Saphra, Naomi |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Automatic VQA Evaluation Using Large Language Models
by: Mañas, Oscar, et al.
Published: (2023)
by: Mañas, Oscar, et al.
Published: (2023)
Library Learning Doesn't: The Curious Case of the Single-Use "Library"
by: Berlot-Attwell, Ian, et al.
Published: (2024)
by: Berlot-Attwell, Ian, et al.
Published: (2024)
Using Shapley interactions to understand how models use structure
by: Singhvi, Divyansh, et al.
Published: (2024)
by: Singhvi, Divyansh, et al.
Published: (2024)
LLM Library Learning Fails: A LEGO-Prover Case Study
by: Berlot-Attwell, Ian, et al.
Published: (2025)
by: Berlot-Attwell, Ian, et al.
Published: (2025)
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
by: Ma, Dongsheng, et al.
Published: (2026)
by: Ma, Dongsheng, et al.
Published: (2026)
BloomVQA: Assessing Hierarchical Multi-modal Comprehension
by: Gong, Yunye, et al.
Published: (2023)
by: Gong, Yunye, et al.
Published: (2023)
No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
by: Udandarao, Vishaal, et al.
Published: (2024)
by: Udandarao, Vishaal, et al.
Published: (2024)
Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations
by: Srivastava, Archita, et al.
Published: (2025)
by: Srivastava, Archita, et al.
Published: (2025)
Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQA
by: Mo, Wentao, et al.
Published: (2024)
by: Mo, Wentao, et al.
Published: (2024)
Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy
by: Ging, Simon, et al.
Published: (2024)
by: Ging, Simon, et al.
Published: (2024)
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
by: Kahl, Kim-Celine, et al.
Published: (2024)
by: Kahl, Kim-Celine, et al.
Published: (2024)
KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
by: Hwang, Taebaek, et al.
Published: (2025)
by: Hwang, Taebaek, et al.
Published: (2025)
GRAM: Global Reasoning for Multi-Page VQA
by: Blau, Tsachi, et al.
Published: (2024)
by: Blau, Tsachi, et al.
Published: (2024)
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts
by: Singh, Shubhankar, et al.
Published: (2024)
by: Singh, Shubhankar, et al.
Published: (2024)
MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning
by: Kumar, Somnath, et al.
Published: (2024)
by: Kumar, Somnath, et al.
Published: (2024)
Systematic Reward Gap Optimization for Mitigating VLM Hallucinations
by: He, Lehan, et al.
Published: (2024)
by: He, Lehan, et al.
Published: (2024)
MAGIC: Near-Optimal Data Attribution for Deep Learning
by: Ilyas, Andrew, et al.
Published: (2025)
by: Ilyas, Andrew, et al.
Published: (2025)
Multimodal LLM With Hierarchical Mixture-of-Experts for VQA on 3D Brain MRI
by: Vepa, Arvind Murari, et al.
Published: (2025)
by: Vepa, Arvind Murari, et al.
Published: (2025)
FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA
by: Sarwar, Nobin
Published: (2025)
by: Sarwar, Nobin
Published: (2025)
NegVQA: Can Vision Language Models Understand Negation?
by: Zhang, Yuhui, et al.
Published: (2025)
by: Zhang, Yuhui, et al.
Published: (2025)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
CommVQA: Situating Visual Question Answering in Communicative Contexts
by: Naik, Nandita Shankar, et al.
Published: (2024)
by: Naik, Nandita Shankar, et al.
Published: (2024)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
by: Ge, Jinchao, et al.
Published: (2025)
by: Ge, Jinchao, et al.
Published: (2025)
Efficient Whole Slide Pathology VQA via Token Compression
by: Lyu, Weimin, et al.
Published: (2025)
by: Lyu, Weimin, et al.
Published: (2025)
Realistic Evaluation of Model Merging for Compositional Generalization
by: Tam, Derek, et al.
Published: (2024)
by: Tam, Derek, et al.
Published: (2024)
Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts
by: Wu, Xuyang, et al.
Published: (2024)
by: Wu, Xuyang, et al.
Published: (2024)
VisMin: Visual Minimal-Change Understanding
by: Awal, Rabiul, et al.
Published: (2024)
by: Awal, Rabiul, et al.
Published: (2024)
Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
by: Ratzlaff, Neale, et al.
Published: (2024)
by: Ratzlaff, Neale, et al.
Published: (2024)
Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval
by: Li, Siting, et al.
Published: (2025)
by: Li, Siting, et al.
Published: (2025)
VLLaVO: Mitigating Visual Gap through LLMs
by: Chen, Shuhao, et al.
Published: (2024)
by: Chen, Shuhao, et al.
Published: (2024)
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
by: Abbasi, Reza, et al.
Published: (2024)
by: Abbasi, Reza, et al.
Published: (2024)
Few-Shot VQA with Frozen LLMs: A Tale of Two Approaches
by: Sterner, Igor, et al.
Published: (2024)
by: Sterner, Igor, et al.
Published: (2024)
FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering
by: Huang, Chengyue, et al.
Published: (2025)
by: Huang, Chengyue, et al.
Published: (2025)
Bridging the Semantic Gaps: Improving Medical VQA Consistency with LLM-Augmented Question Sets
by: Ma, Yongpei, et al.
Published: (2025)
by: Ma, Yongpei, et al.
Published: (2025)
Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry
by: Hou, Wenjun, et al.
Published: (2024)
by: Hou, Wenjun, et al.
Published: (2024)
IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models
by: Shahgir, Haz Sameen, et al.
Published: (2024)
by: Shahgir, Haz Sameen, et al.
Published: (2024)
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
by: Yanuka, Moran, et al.
Published: (2024)
by: Yanuka, Moran, et al.
Published: (2024)
MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning
by: Yu, Suhao, et al.
Published: (2025)
by: Yu, Suhao, et al.
Published: (2025)
Towards Efficient and Robust VQA-NLE Data Generation with Large Vision-Language Models
by: Irawan, Patrick Amadeus, et al.
Published: (2024)
by: Irawan, Patrick Amadeus, et al.
Published: (2024)
Similar Items
-
Improving Automatic VQA Evaluation Using Large Language Models
by: Mañas, Oscar, et al.
Published: (2023) -
Library Learning Doesn't: The Curious Case of the Single-Use "Library"
by: Berlot-Attwell, Ian, et al.
Published: (2024) -
Using Shapley interactions to understand how models use structure
by: Singhvi, Divyansh, et al.
Published: (2024) -
LLM Library Learning Fails: A LEGO-Prover Case Study
by: Berlot-Attwell, Ian, et al.
Published: (2025) -
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
by: Ma, Dongsheng, et al.
Published: (2026)