Unveiling the Visual Counting Bottleneck in Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pang, Xingzhou, Hou, Yifan, Wang, Junling, Sachan, Mrinmaya
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914614457925632
author Pang, Xingzhou
Hou, Yifan
Wang, Junling
Sachan, Mrinmaya
author_facet Pang, Xingzhou
Hou, Yifan
Wang, Junling
Sachan, Mrinmaya
contents While Large Vision-Language Models (VLMs) excel at interpolation, they suffer catastrophic failures in systematic generalization, most notably in visual counting. In this work, we investigate this extrapolation bottleneck by deconstructing visual counting into three cognitive stages: visual individuation, magnitude awareness, and symbolic mapping. Using synthetic Go boards and linear probes, we demonstrate that visual backbones maintain robust, linearly separable representations of quantity well into the extrapolation regime, ruling out perceptual failure. Furthermore, models retain latent magnitude awareness, successfully performing comparative reasoning on quantities they fail to enumerate. We pinpoint the collapse to the symbolic mapping stage, where the model fails to project valid visual magnitudes onto symbolic tokens. Our findings support a frac tured magnitude hypothesis: VLMs fail to acquire a universal number space, instead learning disjoint, modality-specific statistical manifolds that prevent cross-modal grounding for unseen quantities. Validated on the state-of-the-art foundation model, our results suggest that bridging this gap requires inductive priors enforcing unified representations, as data scaling alone is insufficient.
format Preprint
id arxiv_https___arxiv_org_abs_2605_30170
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Unveiling the Visual Counting Bottleneck in Vision-Language Models
Pang, Xingzhou
Hou, Yifan
Wang, Junling
Sachan, Mrinmaya
Multimedia
Computer Vision and Pattern Recognition
Machine Learning
While Large Vision-Language Models (VLMs) excel at interpolation, they suffer catastrophic failures in systematic generalization, most notably in visual counting. In this work, we investigate this extrapolation bottleneck by deconstructing visual counting into three cognitive stages: visual individuation, magnitude awareness, and symbolic mapping. Using synthetic Go boards and linear probes, we demonstrate that visual backbones maintain robust, linearly separable representations of quantity well into the extrapolation regime, ruling out perceptual failure. Furthermore, models retain latent magnitude awareness, successfully performing comparative reasoning on quantities they fail to enumerate. We pinpoint the collapse to the symbolic mapping stage, where the model fails to project valid visual magnitudes onto symbolic tokens. Our findings support a frac tured magnitude hypothesis: VLMs fail to acquire a universal number space, instead learning disjoint, modality-specific statistical manifolds that prevent cross-modal grounding for unseen quantities. Validated on the state-of-the-art foundation model, our results suggest that bridging this gap requires inductive priors enforcing unified representations, as data scaling alone is insufficient.
title Unveiling the Visual Counting Bottleneck in Vision-Language Models
topic Multimedia
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2605.30170