Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Guo, Xuyang, Huang, Zekai, Shi, Zhenmei, Song, Zhao, Zhang, Jiahao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918154743054336
author Guo, Xuyang
Huang, Zekai
Shi, Zhenmei
Song, Zhao
Zhang, Jiahao
author_facet Guo, Xuyang
Huang, Zekai
Shi, Zhenmei
Song, Zhao
Zhang, Jiahao
contents Vision-Language Models (VLMs) have become a central focus of today's AI community, owing to their impressive abilities gained from training on large-scale vision-language data from the Web. These models have demonstrated strong performance across diverse tasks, including image understanding, video understanding, complex visual reasoning, and embodied AI. Despite these noteworthy successes, a fundamental question remains: Can VLMs count objects correctly? In this paper, we introduce a simple yet effective benchmark, VLMCountBench, designed under a minimalist setting with only basic geometric shapes (e.g., triangles, circles) and their compositions, focusing exclusively on counting tasks without interference from other factors. We adopt strict independent variable control and systematically study the effects of simple properties such as color, size, and prompt refinement in a controlled ablation. Our empirical results reveal that while VLMs can count reliably when only one shape type is present, they exhibit substantial failures when multiple shape types are combined (i.e., compositional counting). This highlights a fundamental empirical limitation of current VLMs and motivates important directions for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04401
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
Guo, Xuyang
Huang, Zekai
Shi, Zhenmei
Song, Zhao
Zhang, Jiahao
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision-Language Models (VLMs) have become a central focus of today's AI community, owing to their impressive abilities gained from training on large-scale vision-language data from the Web. These models have demonstrated strong performance across diverse tasks, including image understanding, video understanding, complex visual reasoning, and embodied AI. Despite these noteworthy successes, a fundamental question remains: Can VLMs count objects correctly? In this paper, we introduce a simple yet effective benchmark, VLMCountBench, designed under a minimalist setting with only basic geometric shapes (e.g., triangles, circles) and their compositions, focusing exclusively on counting tasks without interference from other factors. We adopt strict independent variable control and systematically study the effects of simple properties such as color, size, and prompt refinement in a controlled ablation. Our empirical results reveal that while VLMs can count reliably when only one shape type is present, they exhibit substantial failures when multiple shape types are combined (i.e., compositional counting). This highlights a fundamental empirical limitation of current VLMs and motivates important directions for future research.
title Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.04401