Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Aravindan, Ashwath Vaithinathan, Jha, Abha, Kulkarni, Mihir |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2025)
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2025)
Backdoor Defense in Diffusion Models via Spatial Attention Unlearning
von: Jha, Abha, et al.
Veröffentlicht: (2025)
von: Jha, Abha, et al.
Veröffentlicht: (2025)
VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following
von: Hong, Hyesoo, et al.
Veröffentlicht: (2026)
von: Hong, Hyesoo, et al.
Veröffentlicht: (2026)
Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)
CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution
von: Tian, Baoliang, et al.
Veröffentlicht: (2025)
von: Tian, Baoliang, et al.
Veröffentlicht: (2025)
NICE FACT: Diagnosing and Calibrating VLMs in Quantitative Reasoning for Kinematic Physics
von: Lan, Jian, et al.
Veröffentlicht: (2026)
von: Lan, Jian, et al.
Veröffentlicht: (2026)
UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs
von: Ni, Shuo, et al.
Veröffentlicht: (2026)
von: Ni, Shuo, et al.
Veröffentlicht: (2026)
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
von: Li, Shuo, et al.
Veröffentlicht: (2024)
von: Li, Shuo, et al.
Veröffentlicht: (2024)
Eye Gaze Tells You Where to Compute: Gaze-Driven Efficient VLMs
von: Chen, Qinyu, et al.
Veröffentlicht: (2025)
von: Chen, Qinyu, et al.
Veröffentlicht: (2025)
VLMs Guided Interpretable Decision Making for Autonomous Driving
von: Hu, Xin, et al.
Veröffentlicht: (2025)
von: Hu, Xin, et al.
Veröffentlicht: (2025)
CLUENet: Cluster Attention Makes Neural Networks Have Eyes
von: Song, Xiangshuai, et al.
Veröffentlicht: (2025)
von: Song, Xiangshuai, et al.
Veröffentlicht: (2025)
Smart Eyes for Silent Threats: VLMs and In-Context Learning for THz Imaging
von: Poggi, Nicolas, et al.
Veröffentlicht: (2025)
von: Poggi, Nicolas, et al.
Veröffentlicht: (2025)
Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols
von: Zeng, Xianchao, et al.
Veröffentlicht: (2025)
von: Zeng, Xianchao, et al.
Veröffentlicht: (2025)
MedConcept: Unsupervised Concept Discovery for Interpretability in Medical VLMs
von: Haque, Md Rakibul, et al.
Veröffentlicht: (2026)
von: Haque, Md Rakibul, et al.
Veröffentlicht: (2026)
Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
Multi-View Camera System for Variant-Aware Autonomous Vehicle Inspection and Defect Detection
von: Kulkarni, Yash, et al.
Veröffentlicht: (2025)
von: Kulkarni, Yash, et al.
Veröffentlicht: (2025)
Interpret, prune and distill Donut : towards lightweight VLMs for VQA on document
von: Mansour, Adnan Ben, et al.
Veröffentlicht: (2025)
von: Mansour, Adnan Ben, et al.
Veröffentlicht: (2025)
ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
von: Huang, Irene, et al.
Veröffentlicht: (2024)
von: Huang, Irene, et al.
Veröffentlicht: (2024)
DISSECT: Diagnosing Where Vision Ends and Language Priors Begin in Scientific VLMs
von: Kukreja, Dikshant, et al.
Veröffentlicht: (2026)
von: Kukreja, Dikshant, et al.
Veröffentlicht: (2026)
BioVLM: Routing Prompts, Not Parameters, for Cross-Modality Generalization in Biomedical VLMs
von: Singha, Mainak, et al.
Veröffentlicht: (2026)
von: Singha, Mainak, et al.
Veröffentlicht: (2026)
Explainable Transformer Prototypes for Medical Diagnoses
von: Demir, Ugur, et al.
Veröffentlicht: (2024)
von: Demir, Ugur, et al.
Veröffentlicht: (2024)
Continuous Perception Matters: Diagnosing Temporal Integration Failures in Multimodal Models
von: Wang, Zeyu, et al.
Veröffentlicht: (2024)
von: Wang, Zeyu, et al.
Veröffentlicht: (2024)
Rewarding Intellectual Humility Learning When Not To Answer In Large Language Models
von: Jha, Abha, et al.
Veröffentlicht: (2026)
von: Jha, Abha, et al.
Veröffentlicht: (2026)
Dissecting and Mitigating Diffusion Bias via Mechanistic Interpretability
von: Shi, Yingdong, et al.
Veröffentlicht: (2025)
von: Shi, Yingdong, et al.
Veröffentlicht: (2025)
GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMs
von: Mantini, Pranav, et al.
Veröffentlicht: (2026)
von: Mantini, Pranav, et al.
Veröffentlicht: (2026)
Robust Object Detection with Pseudo Labels from VLMs using Per-Object Co-teaching
von: Bhaskar, Uday, et al.
Veröffentlicht: (2025)
von: Bhaskar, Uday, et al.
Veröffentlicht: (2025)
Do VLMs Have a Moral Backbone? A Study on the Fragile Morality of Vision-Language Models
von: Liu, Zhining, et al.
Veröffentlicht: (2026)
von: Liu, Zhining, et al.
Veröffentlicht: (2026)
Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
von: Li, Yiwei, et al.
Veröffentlicht: (2026)
von: Li, Yiwei, et al.
Veröffentlicht: (2026)
Interpreting Biomedical VLMs on High-Imbalance Out-of-Distributions: An Insight into BiomedCLIP on Radiology
von: Sadman, Nafiz, et al.
Veröffentlicht: (2025)
von: Sadman, Nafiz, et al.
Veröffentlicht: (2025)
Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning
von: Xie, Zhuofan, et al.
Veröffentlicht: (2026)
von: Xie, Zhuofan, et al.
Veröffentlicht: (2026)
Evaluating Compositional Generalisation in VLMs and Diffusion Models
von: Pearson, Beth, et al.
Veröffentlicht: (2025)
von: Pearson, Beth, et al.
Veröffentlicht: (2025)
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs
von: Salazar, Israfel, et al.
Veröffentlicht: (2025)
von: Salazar, Israfel, et al.
Veröffentlicht: (2025)
Scale Alone Does not Improve Mechanistic Interpretability in Vision Models
von: Zimmermann, Roland S., et al.
Veröffentlicht: (2023)
von: Zimmermann, Roland S., et al.
Veröffentlicht: (2023)
Synthetic Designed Experiments for Diagnosing Vision Model Failure
von: Sarkar, Krisanu
Veröffentlicht: (2026)
von: Sarkar, Krisanu
Veröffentlicht: (2026)
Interpretable Failure Detection with Human-Level Concepts
von: Nguyen, Kien X., et al.
Veröffentlicht: (2025)
von: Nguyen, Kien X., et al.
Veröffentlicht: (2025)
PRIME: Prioritizing Interpretability in Failure Mode Extraction
von: Rezaei, Keivan, et al.
Veröffentlicht: (2023)
von: Rezaei, Keivan, et al.
Veröffentlicht: (2023)
Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusions
von: Sun, Xiaoxiao, et al.
Veröffentlicht: (2026)
von: Sun, Xiaoxiao, et al.
Veröffentlicht: (2026)
Manipulating a Tetris-Inspired 3D Video Representation
von: Godbole, Mihir
Veröffentlicht: (2024)
von: Godbole, Mihir
Veröffentlicht: (2024)
Rethinking Visual Privacy: A Compositional Privacy Risk Framework for Severity Assessment with VLMs
von: Tsaprazlis, Efthymios, et al.
Veröffentlicht: (2026)
von: Tsaprazlis, Efthymios, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2025) -
Backdoor Defense in Diffusion Models via Spatial Attention Unlearning
von: Jha, Abha, et al.
Veröffentlicht: (2025) -
VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following
von: Hong, Hyesoo, et al.
Veröffentlicht: (2026) -
Second Guess: Detecting Uncertainty Through Abstention and Answer Stability in Small Language Models
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026) -
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)