VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gavrikov, Paul, Lin, Wei, Mirza, M. Jehanzeb, Jahagirdar, Soumya, Huzaifa, Muhammad, Doveh, Sivan, Yeung-Levy, Serena, Glass, James, Kuehne, Hilde
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913158538461184
author Gavrikov, Paul
Lin, Wei
Mirza, M. Jehanzeb
Jahagirdar, Soumya
Huzaifa, Muhammad
Doveh, Sivan
Yeung-Levy, Serena
Glass, James
Kuehne, Hilde
author_facet Gavrikov, Paul
Lin, Wei
Mirza, M. Jehanzeb
Jahagirdar, Soumya
Huzaifa, Muhammad
Doveh, Sivan
Yeung-Levy, Serena
Glass, James
Kuehne, Hilde
contents Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answering (VQA) benchmark comprising 2,720 question-answer pairs, with privately held ground-truth responses. Unlike prior VQA datasets that typically focus on near global image understanding, VisualOverload challenges models to perform simple, knowledge-free vision tasks in densely populated (or, overloaded) scenes. Our dataset consists of high-resolution scans of public-domain paintings that are populated with multiple figures, actions, and unfolding subplots set against elaborately detailed backdrops. We manually annotated these images with questions across six task categories to probe for a thorough understanding of the scene. We hypothesize that current benchmarks overestimate the performance of VLMs, and encoding and reasoning over details is still a challenging task for them, especially if they are confronted with densely populated scenes. Indeed, we observe that even the best model (o3) out of 37 tested models only achieves 19.6% accuracy on our hardest test split and overall 69.5% accuracy on all questions. Beyond a thorough evaluation, we complement our benchmark with an error analysis that reveals multiple failure modes, including a lack of counting skills, failure in OCR, and striking logical inconsistencies under complex tasks. Altogether, VisualOverload exposes a critical gap in current vision models and offers a crucial resource for the community to develop better models. Benchmark: http://paulgavrikov.github.io/visualoverload
format Preprint
id arxiv_https___arxiv_org_abs_2509_25339
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
Gavrikov, Paul
Lin, Wei
Mirza, M. Jehanzeb
Jahagirdar, Soumya
Huzaifa, Muhammad
Doveh, Sivan
Yeung-Levy, Serena
Glass, James
Kuehne, Hilde
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Image and Video Processing
Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answering (VQA) benchmark comprising 2,720 question-answer pairs, with privately held ground-truth responses. Unlike prior VQA datasets that typically focus on near global image understanding, VisualOverload challenges models to perform simple, knowledge-free vision tasks in densely populated (or, overloaded) scenes. Our dataset consists of high-resolution scans of public-domain paintings that are populated with multiple figures, actions, and unfolding subplots set against elaborately detailed backdrops. We manually annotated these images with questions across six task categories to probe for a thorough understanding of the scene. We hypothesize that current benchmarks overestimate the performance of VLMs, and encoding and reasoning over details is still a challenging task for them, especially if they are confronted with densely populated scenes. Indeed, we observe that even the best model (o3) out of 37 tested models only achieves 19.6% accuracy on our hardest test split and overall 69.5% accuracy on all questions. Beyond a thorough evaluation, we complement our benchmark with an error analysis that reveals multiple failure modes, including a lack of counting skills, failure in OCR, and striking logical inconsistencies under complex tasks. Altogether, VisualOverload exposes a critical gap in current vision models and offers a crucial resource for the community to develop better models. Benchmark: http://paulgavrikov.github.io/visualoverload
title VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Image and Video Processing
url https://arxiv.org/abs/2509.25339