Edge Reliability Gap in Vision-Language Models: Quantifying Failure Modes of Compressed VLMs Under Visual Corruption
Fuente:
arXiv
Salvato in:
| Autore principale: | Erol, Mehmet Kaan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Discovering Failure Modes in Vision-Language Models using RL
di: Jain, Kanishk, et al.
Pubblicazione: (2026)
di: Jain, Kanishk, et al.
Pubblicazione: (2026)
Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
Towards Lossless Ultimate Vision Token Compression for VLMs
di: Zheng, Dehua, et al.
Pubblicazione: (2025)
di: Zheng, Dehua, et al.
Pubblicazione: (2025)
Mixed Signals: Decoding VLMs' Reasoning and Underlying Bias in Vision-Language Conflict
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2025)
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2025)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
di: Berman, Shmuel, et al.
Pubblicazione: (2025)
di: Berman, Shmuel, et al.
Pubblicazione: (2025)
AMVICC: A Novel Benchmark for Cross-Modal Failure Mode Profiling for VLMs and IGMs
di: Basappa, Aahana, et al.
Pubblicazione: (2026)
di: Basappa, Aahana, et al.
Pubblicazione: (2026)
VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following
di: Hong, Hyesoo, et al.
Pubblicazione: (2026)
di: Hong, Hyesoo, et al.
Pubblicazione: (2026)
SycoPhantasy: Quantifying Sycophancy and Hallucination in Small Open Weight VLMs for Vision-Language Scoring of Fantasy Characters
di: Shah, Arya, et al.
Pubblicazione: (2026)
di: Shah, Arya, et al.
Pubblicazione: (2026)
Evaluating Vision Language Models (VLMs) for Radiology: A Comprehensive Analysis
di: Li, Frank, et al.
Pubblicazione: (2025)
di: Li, Frank, et al.
Pubblicazione: (2025)
Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables
di: Singh, Anshul, et al.
Pubblicazione: (2025)
di: Singh, Anshul, et al.
Pubblicazione: (2025)
Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
di: Wang, Huanyu, et al.
Pubblicazione: (2025)
di: Wang, Huanyu, et al.
Pubblicazione: (2025)
ViFP: A Framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMs
di: Zhang, Ben, et al.
Pubblicazione: (2025)
di: Zhang, Ben, et al.
Pubblicazione: (2025)
The Geometry of Representational Failures in Vision Language Models
di: Savietto, Daniele, et al.
Pubblicazione: (2026)
di: Savietto, Daniele, et al.
Pubblicazione: (2026)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
di: Zhou, Guanyu, et al.
Pubblicazione: (2026)
di: Zhou, Guanyu, et al.
Pubblicazione: (2026)
Cross-Modal Attention Analysis and Optimization in Vision-Language Models: A Study on Visual Reliability
di: Zhou, Lijie
Pubblicazione: (2026)
di: Zhou, Lijie
Pubblicazione: (2026)
Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning
di: Zhang, Gengyuan, et al.
Pubblicazione: (2023)
di: Zhang, Gengyuan, et al.
Pubblicazione: (2023)
DISSECT: Diagnosing Where Vision Ends and Language Priors Begin in Scientific VLMs
di: Kukreja, Dikshant, et al.
Pubblicazione: (2026)
di: Kukreja, Dikshant, et al.
Pubblicazione: (2026)
NanoVLMs: How small can we go and still make coherent Vision Language Models?
di: Agarwalla, Mukund, et al.
Pubblicazione: (2025)
di: Agarwalla, Mukund, et al.
Pubblicazione: (2025)
FCoT-VL:Advancing Text-oriented Large Vision-Language Models with Efficient Visual Token Compression
di: Li, Jianjian, et al.
Pubblicazione: (2025)
di: Li, Jianjian, et al.
Pubblicazione: (2025)
Collaborative Edge-to-Server Inference for Vision-Language Models
di: Song, Soochang, et al.
Pubblicazione: (2025)
di: Song, Soochang, et al.
Pubblicazione: (2025)
VACoT: Rethinking Visual Data Augmentation with VLMs
di: Xu, Zhengzhuo, et al.
Pubblicazione: (2025)
di: Xu, Zhengzhuo, et al.
Pubblicazione: (2025)
ButterflyViT: 354$\times$ Expert Compression for Edge Vision Transformers
di: Karmore, Aryan
Pubblicazione: (2026)
di: Karmore, Aryan
Pubblicazione: (2026)
Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)
di: Duan, Yicheng, et al.
Pubblicazione: (2025)
di: Duan, Yicheng, et al.
Pubblicazione: (2025)
Evaluating the Impact of Compression Techniques on the Robustness of CNNs under Natural Corruptions
di: Da Silva, Itallo Patrick Castro Alves, et al.
Pubblicazione: (2025)
di: Da Silva, Itallo Patrick Castro Alves, et al.
Pubblicazione: (2025)
MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning
di: Pan, Jiazhen, et al.
Pubblicazione: (2025)
di: Pan, Jiazhen, et al.
Pubblicazione: (2025)
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
di: Kumar, Sunil, et al.
Pubblicazione: (2025)
di: Kumar, Sunil, et al.
Pubblicazione: (2025)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
Can Generalist Vision Language Models (VLMs) Rival Specialist Medical VLMs? Benchmarking and Strategic Insights
di: Zhong, Yuan, et al.
Pubblicazione: (2025)
di: Zhong, Yuan, et al.
Pubblicazione: (2025)
On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression
di: Zhang, Xinwei, et al.
Pubblicazione: (2026)
di: Zhang, Xinwei, et al.
Pubblicazione: (2026)
Aligned Vector Quantization for Edge-Cloud Collabrative Vision-Language Models
di: Liu, Xiao, et al.
Pubblicazione: (2024)
di: Liu, Xiao, et al.
Pubblicazione: (2024)
KPCA-CAM: Visual Explainability of Deep Computer Vision Models using Kernel PCA
di: Karmani, Sachin, et al.
Pubblicazione: (2024)
di: Karmani, Sachin, et al.
Pubblicazione: (2024)
LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs
di: Luo, Kun, et al.
Pubblicazione: (2026)
di: Luo, Kun, et al.
Pubblicazione: (2026)
HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
di: Zhang, Jusheng, et al.
Pubblicazione: (2025)
di: Zhang, Jusheng, et al.
Pubblicazione: (2025)
To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs
di: Hong, Rui, et al.
Pubblicazione: (2026)
di: Hong, Rui, et al.
Pubblicazione: (2026)
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
di: Zhang, Qizhe, et al.
Pubblicazione: (2024)
di: Zhang, Qizhe, et al.
Pubblicazione: (2024)
Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling
di: Pantazopoulos, Georgios, et al.
Pubblicazione: (2024)
di: Pantazopoulos, Georgios, et al.
Pubblicazione: (2024)
Skill-Conditioned Visual Geolocation for Vision-Language Models
di: Yang, Chenjie, et al.
Pubblicazione: (2026)
di: Yang, Chenjie, et al.
Pubblicazione: (2026)
Generative Visual Communication in the Era of Vision-Language Models
di: Vinker, Yael
Pubblicazione: (2024)
di: Vinker, Yael
Pubblicazione: (2024)
Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
di: Karamcheti, Siddharth, et al.
Pubblicazione: (2024)
di: Karamcheti, Siddharth, et al.
Pubblicazione: (2024)
LL-ICM: Image Compression for Low-level Machine Vision via Large Vision-Language Model
di: Xue, Yuan, et al.
Pubblicazione: (2024)
di: Xue, Yuan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Discovering Failure Modes in Vision-Language Models using RL
di: Jain, Kanishk, et al.
Pubblicazione: (2026) -
Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
di: Guo, Xuyang, et al.
Pubblicazione: (2025) -
Towards Lossless Ultimate Vision Token Compression for VLMs
di: Zheng, Dehua, et al.
Pubblicazione: (2025) -
Mixed Signals: Decoding VLMs' Reasoning and Underlying Bias in Vision-Language Conflict
di: Pezeshkpour, Pouya, et al.
Pubblicazione: (2025) -
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
di: Berman, Shmuel, et al.
Pubblicazione: (2025)