Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qiu, Jason, Meurer, Zachary, Thomas, Xavier, Ghadiyaram, Deepti
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908934271401984
author Qiu, Jason
Meurer, Zachary
Thomas, Xavier
Ghadiyaram, Deepti
author_facet Qiu, Jason
Meurer, Zachary
Thomas, Xavier
Ghadiyaram, Deepti
contents This work investigates the fundamental fragility of state-of-the-art Vision-Language Models (VLMs) under basic geometric transformations. While modern VLMs excel at semantic tasks such as recognizing objects in canonical orientations and describing complex scenes, they exhibit systematic failures at a more fundamental level: lack of robust spatial invariance and equivariance required to reliably determine object identity under simple rotations, scaling, and identity transformations. We demonstrate this limitation through a systematic evaluation across diverse visual domains, including symbolic sketches, natural photographs, and abstract art. Performance drops sharply as semantic content becomes sparse, and this behavior is observed across architectures, model capacities, and prompting strategies. Overall, our results reveal a systematic gap between semantic understanding and spatial reasoning in current VLMs, highlighting the need for stronger geometric grounding in future multimodal systems.
format Preprint
id arxiv_https___arxiv_org_abs_2604_01848
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance
Qiu, Jason
Meurer, Zachary
Thomas, Xavier
Ghadiyaram, Deepti
Computer Vision and Pattern Recognition
This work investigates the fundamental fragility of state-of-the-art Vision-Language Models (VLMs) under basic geometric transformations. While modern VLMs excel at semantic tasks such as recognizing objects in canonical orientations and describing complex scenes, they exhibit systematic failures at a more fundamental level: lack of robust spatial invariance and equivariance required to reliably determine object identity under simple rotations, scaling, and identity transformations. We demonstrate this limitation through a systematic evaluation across diverse visual domains, including symbolic sketches, natural photographs, and abstract art. Performance drops sharply as semantic content becomes sparse, and this behavior is observed across architectures, model capacities, and prompting strategies. Overall, our results reveal a systematic gap between semantic understanding and spatial reasoning in current VLMs, highlighting the need for stronger geometric grounding in future multimodal systems.
title Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.01848