Right Predictions, Misleading Explanations: On the Vulnerability of Vision-Language Model Explanations
Fuente:
arXiv
Saved in:
| Main Authors: | Babadi, Narges, Karimipour, Hadis |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data-centric Prediction Explanation via Kernelized Stein Discrepancy
by: Sarvmaili, Mahtab, et al.
Published: (2024)
by: Sarvmaili, Mahtab, et al.
Published: (2024)
How to Squeeze An Explanation Out of Your Model
by: Roxo, Tiago, et al.
Published: (2024)
by: Roxo, Tiago, et al.
Published: (2024)
LINE: LLM-based Iterative Neuron Explanations for Vision Models
by: Zaigrajew, Vladimir, et al.
Published: (2026)
by: Zaigrajew, Vladimir, et al.
Published: (2026)
Explanation Bottleneck Models
by: Yamaguchi, Shin'ya, et al.
Published: (2024)
by: Yamaguchi, Shin'ya, et al.
Published: (2024)
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
by: Lagos, Maximiliano Hormazábal, et al.
Published: (2025)
by: Lagos, Maximiliano Hormazábal, et al.
Published: (2025)
FovEx: Human-Inspired Explanations for Vision Transformers and Convolutional Neural Networks
by: Panda, Mahadev Prasad, et al.
Published: (2024)
by: Panda, Mahadev Prasad, et al.
Published: (2024)
Activation Matching for Explanation Generation
by: Suhail, Pirzada, et al.
Published: (2025)
by: Suhail, Pirzada, et al.
Published: (2025)
Linear Explanations for Individual Neurons
by: Oikarinen, Tuomas, et al.
Published: (2024)
by: Oikarinen, Tuomas, et al.
Published: (2024)
ViGText: Deepfake Image Detection with Vision-Language Model Explanations and Graph Neural Networks
by: ALBarqawi, Ahmad, et al.
Published: (2025)
by: ALBarqawi, Ahmad, et al.
Published: (2025)
Probabilistic Conceptual Explainers: Trustworthy Conceptual Explanations for Vision Foundation Models
by: Wang, Hengyi, et al.
Published: (2024)
by: Wang, Hengyi, et al.
Published: (2024)
Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions
by: Indrehus, Kjetil, et al.
Published: (2026)
by: Indrehus, Kjetil, et al.
Published: (2026)
TACE: Tumor-Aware Counterfactual Explanations
by: Rossi, Eleonora Beatrice, et al.
Published: (2024)
by: Rossi, Eleonora Beatrice, et al.
Published: (2024)
Applying Graph Explanation to Operator Fusion
by: Mills, Keith G., et al.
Published: (2024)
by: Mills, Keith G., et al.
Published: (2024)
The Manifold Hypothesis for Gradient-Based Explanations
by: Bordt, Sebastian, et al.
Published: (2022)
by: Bordt, Sebastian, et al.
Published: (2022)
Revealing Vulnerabilities of Neural Networks in Parameter Learning and Defense Against Explanation-Aware Backdoors
by: Kadir, Md Abdul, et al.
Published: (2024)
by: Kadir, Md Abdul, et al.
Published: (2024)
What Helps---and What Hurts: Bidirectional Explanations for Vision Transformers
by: Su, Qin, et al.
Published: (2026)
by: Su, Qin, et al.
Published: (2026)
Robustness of Visual Explanations to Common Data Augmentation
by: Tětková, Lenka, et al.
Published: (2023)
by: Tětková, Lenka, et al.
Published: (2023)
On the Generalization and Causal Explanation in Self-Supervised Learning
by: Qiang, Wenwen, et al.
Published: (2024)
by: Qiang, Wenwen, et al.
Published: (2024)
LD-ViCE: Latent Diffusion Model for Video Counterfactual Explanations
by: Varshney, Payal, et al.
Published: (2025)
by: Varshney, Payal, et al.
Published: (2025)
Cross-Task Affinity Learning for Multitask Dense Scene Predictions
by: Sinodinos, Dimitrios, et al.
Published: (2024)
by: Sinodinos, Dimitrios, et al.
Published: (2024)
Why Does It Look There? Structured Explanations for Image Classification
by: Li, Jiarui, et al.
Published: (2026)
by: Li, Jiarui, et al.
Published: (2026)
Architecture-Aware Explanation Auditing for Industrial Visual Inspection
by: Jia, Sibo, et al.
Published: (2026)
by: Jia, Sibo, et al.
Published: (2026)
Discovering and Mitigating Visual Biases through Keyword Explanation
by: Kim, Younghyun, et al.
Published: (2023)
by: Kim, Younghyun, et al.
Published: (2023)
Disentangled Explanations of Neural Network Predictions by Finding Relevant Subspaces
by: Chormai, Pattarawat, et al.
Published: (2022)
by: Chormai, Pattarawat, et al.
Published: (2022)
Non-identifiability of Explanations from Model Behavior in Deep Networks of Image Authenticity Judgments
by: Depaolini, Icaro Re, et al.
Published: (2026)
by: Depaolini, Icaro Re, et al.
Published: (2026)
SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations
by: Chen, Zhaorun, et al.
Published: (2024)
by: Chen, Zhaorun, et al.
Published: (2024)
Which Models have Perceptually-Aligned Gradients? An Explanation via Off-Manifold Robustness
by: Srinivas, Suraj, et al.
Published: (2023)
by: Srinivas, Suraj, et al.
Published: (2023)
xMIL: Insightful Explanations for Multiple Instance Learning in Histopathology
by: Hense, Julius, et al.
Published: (2024)
by: Hense, Julius, et al.
Published: (2024)
VERITAS: Verification and Explanation of Realness in Images for Transparency in AI Systems
by: Srivastava, Aadi, et al.
Published: (2025)
by: Srivastava, Aadi, et al.
Published: (2025)
Studying How to Efficiently and Effectively Guide Models with Explanations
by: Rao, Sukrut, et al.
Published: (2023)
by: Rao, Sukrut, et al.
Published: (2023)
Enhancing Adversarial Example Detection Through Model Explanation
by: Ma, Qian, et al.
Published: (2025)
by: Ma, Qian, et al.
Published: (2025)
Towards a Mechanistic Explanation of Diffusion Model Generalization
by: Niedoba, Matthew, et al.
Published: (2024)
by: Niedoba, Matthew, et al.
Published: (2024)
Beyond Attention Heatmaps: How to Get Better Explanations for Multiple Instance Learning Models in Histopathology
by: Idaji, Mina Jamshidi, et al.
Published: (2026)
by: Idaji, Mina Jamshidi, et al.
Published: (2026)
The Gaussian Discriminant Variational Autoencoder (GdVAE): A Self-Explainable Model with Counterfactual Explanations
by: Haselhoff, Anselm, et al.
Published: (2024)
by: Haselhoff, Anselm, et al.
Published: (2024)
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
ADVISE: ADaptive Feature Relevance and VISual Explanations for Convolutional Neural Networks
by: Dehshibi, Mohammad Mahdi, et al.
Published: (2022)
by: Dehshibi, Mohammad Mahdi, et al.
Published: (2022)
Explanation-based Training with Differentiable Insertion/Deletion Metric-aware Regularizers
by: Yoshikawa, Yuya, et al.
Published: (2023)
by: Yoshikawa, Yuya, et al.
Published: (2023)
Counterfactual Explanations for Medical Image Classification and Regression using Diffusion Autoencoder
by: Atad, Matan, et al.
Published: (2024)
by: Atad, Matan, et al.
Published: (2024)
Explaining Deep Convolutional Neural Networks for Image Classification by Evolving Local Interpretable Model-agnostic Explanations
by: Wang, Bin, et al.
Published: (2022)
by: Wang, Bin, et al.
Published: (2022)
MEGL: Multimodal Explanation-Guided Learning
by: Zhang, Yifei, et al.
Published: (2024)
by: Zhang, Yifei, et al.
Published: (2024)
Similar Items
-
Data-centric Prediction Explanation via Kernelized Stein Discrepancy
by: Sarvmaili, Mahtab, et al.
Published: (2024) -
How to Squeeze An Explanation Out of Your Model
by: Roxo, Tiago, et al.
Published: (2024) -
LINE: LLM-based Iterative Neuron Explanations for Vision Models
by: Zaigrajew, Vladimir, et al.
Published: (2026) -
Explanation Bottleneck Models
by: Yamaguchi, Shin'ya, et al.
Published: (2024) -
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
by: Lagos, Maximiliano Hormazábal, et al.
Published: (2025)