The Geometry of Representational Failures in Vision Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Savietto, Daniele, Campbell, Declan, Panisson, André, Nurisso, Marco, Petri, Giovanni, Cohen, Jonathan D., Perotti, Alan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders
by: Cassano, Enrico, et al.
Published: (2025)
by: Cassano, Enrico, et al.
Published: (2025)
HOLMES: HOLonym-MEronym based Semantic inspection for Convolutional Image Classifiers
by: Dibitonto, Francesco, et al.
Published: (2024)
by: Dibitonto, Francesco, et al.
Published: (2024)
Binding Visual Features Point by Point
by: Haputhanthri, Udith, et al.
Published: (2026)
by: Haputhanthri, Udith, et al.
Published: (2026)
Attributes Shape the Embedding Space of Face Recognition Models
by: Leroy, Pierrick, et al.
Published: (2025)
by: Leroy, Pierrick, et al.
Published: (2025)
Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem
by: Campbell, Declan, et al.
Published: (2024)
by: Campbell, Declan, et al.
Published: (2024)
Global Geometry Is Not Enough for Vision Representations
by: Chung, Jiwan, et al.
Published: (2026)
by: Chung, Jiwan, et al.
Published: (2026)
GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation
by: Yang, Jiahao, et al.
Published: (2026)
by: Yang, Jiahao, et al.
Published: (2026)
Discovering Failure Modes in Vision-Language Models using RL
by: Jain, Kanishk, et al.
Published: (2026)
by: Jain, Kanishk, et al.
Published: (2026)
SpatialFly: Geometry-Guided Representation Alignment for UAV Vision-and-Language Navigation in Urban Environments
by: Jiang, Wen, et al.
Published: (2026)
by: Jiang, Wen, et al.
Published: (2026)
GeoWorld-VLM: Geometry from World Models for Vision-Language Models
by: Gu, Renjie, et al.
Published: (2026)
by: Gu, Renjie, et al.
Published: (2026)
GTMA: Dynamic Representation Optimization for OOD Vision-Language Models
by: Zhang, Jensen, et al.
Published: (2025)
by: Zhang, Jensen, et al.
Published: (2025)
REOrdering Patches Improves Vision Models
by: Kutscher, Declan, et al.
Published: (2025)
by: Kutscher, Declan, et al.
Published: (2025)
Proximal Vision Transformer: Enhancing Feature Representation through Two-Stage Manifold Geometry
by: Yun, Haoyu, et al.
Published: (2025)
by: Yun, Haoyu, et al.
Published: (2025)
Frustratingly Easy Test-Time Adaptation of Vision-Language Models
by: Farina, Matteo, et al.
Published: (2024)
by: Farina, Matteo, et al.
Published: (2024)
Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
by: Guo, Xuyang, et al.
Published: (2025)
by: Guo, Xuyang, et al.
Published: (2025)
Seeing the Abstract: Translating the Abstract Language for Vision Language Models
by: Talon, Davide, et al.
Published: (2025)
by: Talon, Davide, et al.
Published: (2025)
Edge Reliability Gap in Vision-Language Models: Quantifying Failure Modes of Compressed VLMs Under Visual Corruption
by: Erol, Mehmet Kaan
Published: (2026)
by: Erol, Mehmet Kaan
Published: (2026)
Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
HarmoCLIP: Harmonizing Global and Regional Representations in Contrastive Vision-Language Models
by: Zeng, Haoxi, et al.
Published: (2025)
by: Zeng, Haoxi, et al.
Published: (2025)
Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning
by: Bleeker, Maurits, et al.
Published: (2024)
by: Bleeker, Maurits, et al.
Published: (2024)
Representation Calibration and Uncertainty Guidance for Class-Incremental Learning based on Vision Language Model
by: Tan, Jiantao, et al.
Published: (2025)
by: Tan, Jiantao, et al.
Published: (2025)
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model
by: Mei, Xiaodong, et al.
Published: (2026)
by: Mei, Xiaodong, et al.
Published: (2026)
Any-to-Any Vision-Language Model for Multimodal X-ray Imaging and Radiological Report Generation
by: Molino, Daniele, et al.
Published: (2025)
by: Molino, Daniele, et al.
Published: (2025)
Text-to-CT Generation via 3D Latent Diffusion Model with Contrastive Vision-Language Pretraining
by: Molino, Daniele, et al.
Published: (2025)
by: Molino, Daniele, et al.
Published: (2025)
CARPE: Context-Aware Image Representation Prioritization via Ensemble for Large Vision-Language Models
by: Lee, Donghee, et al.
Published: (2026)
by: Lee, Donghee, et al.
Published: (2026)
PISA-Bench: The PISA Index as a Multilingual and Multimodal Metric for the Evaluation of Vision-Language Models
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
GeoVLMath: Enhancing Geometry Reasoning in Vision-Language Models via Cross-Modal Reward for Auxiliary Line Creation
by: Guo, Shasha, et al.
Published: (2025)
by: Guo, Shasha, et al.
Published: (2025)
GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models
by: Guan, Yaohan, et al.
Published: (2026)
by: Guan, Yaohan, et al.
Published: (2026)
Dual-Domain Representation Alignment: Bridging 2D and 3D Vision via Geometry-Aware Architecture Search
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
Vision-Language Models Provide Promptable Representations for Reinforcement Learning
by: Chen, William, et al.
Published: (2024)
by: Chen, William, et al.
Published: (2024)
Tinted Frames: Question Framing Blinds Vision-Language Models
by: Fan, Wan-Cyuan, et al.
Published: (2026)
by: Fan, Wan-Cyuan, et al.
Published: (2026)
Human-Like Coarse Object Representations in Vision Models
by: Gizdov, Andrey, et al.
Published: (2026)
by: Gizdov, Andrey, et al.
Published: (2026)
Enhancing Vision-Language Models with Scene Graphs for Traffic Accident Understanding
by: Lohner, Aaron, et al.
Published: (2024)
by: Lohner, Aaron, et al.
Published: (2024)
A Brief Survey on Leveraging Large Scale Vision Models for Enhanced Robot Grasping
by: Kamboj, Abhi, et al.
Published: (2024)
by: Kamboj, Abhi, et al.
Published: (2024)
Is Geometry Enough? An Evaluation of Landmark-Based Gaze Estimation
by: Agostinelli, Daniele, et al.
Published: (2026)
by: Agostinelli, Daniele, et al.
Published: (2026)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
by: Zhang, Jianke, et al.
Published: (2026)
by: Zhang, Jianke, et al.
Published: (2026)
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
by: Gao, Chongkai, et al.
Published: (2025)
by: Gao, Chongkai, et al.
Published: (2025)
Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
by: Wu, Haoyu, et al.
Published: (2025)
by: Wu, Haoyu, et al.
Published: (2025)
From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration
by: Song, Mingyang, et al.
Published: (2025)
by: Song, Mingyang, et al.
Published: (2025)
Bootstrapping Action-Grounded Visual Dynamics in Unified Vision-Language Models
by: Qiu, Yifu, et al.
Published: (2025)
by: Qiu, Yifu, et al.
Published: (2025)
Similar Items
-
SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders
by: Cassano, Enrico, et al.
Published: (2025) -
HOLMES: HOLonym-MEronym based Semantic inspection for Convolutional Image Classifiers
by: Dibitonto, Francesco, et al.
Published: (2024) -
Binding Visual Features Point by Point
by: Haputhanthri, Udith, et al.
Published: (2026) -
Attributes Shape the Embedding Space of Face Recognition Models
by: Leroy, Pierrick, et al.
Published: (2025) -
Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem
by: Campbell, Declan, et al.
Published: (2024)