Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autore principale: Goldshmidt, Roni
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909532396978176
author Goldshmidt, Roni
author_facet Goldshmidt, Roni
contents Interpretability in Vision-Language Models (VLMs) is crucial for trust, debugging, and decision-making in high-stakes applications. We introduce PixelSHAP, a model-agnostic framework extending Shapley-based analysis to structured visual entities. Unlike previous methods focusing on text prompts, PixelSHAP applies to vision-based reasoning by systematically perturbing image objects and quantifying their influence on a VLM's response. PixelSHAP requires no model internals, operating solely on input-output pairs, making it compatible with open-source and commercial models. It supports diverse embedding-based similarity metrics and scales efficiently using optimization techniques inspired by Shapley-based methods. We validate PixelSHAP in autonomous driving, highlighting its ability to enhance interpretability. Key challenges include segmentation sensitivity and object occlusion. Our open-source implementation facilitates further research.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06670
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On
Goldshmidt, Roni
Computer Vision and Pattern Recognition
Computation and Language
Interpretability in Vision-Language Models (VLMs) is crucial for trust, debugging, and decision-making in high-stakes applications. We introduce PixelSHAP, a model-agnostic framework extending Shapley-based analysis to structured visual entities. Unlike previous methods focusing on text prompts, PixelSHAP applies to vision-based reasoning by systematically perturbing image objects and quantifying their influence on a VLM's response. PixelSHAP requires no model internals, operating solely on input-output pairs, making it compatible with open-source and commercial models. It supports diverse embedding-based similarity metrics and scales efficiently using optimization techniques inspired by Shapley-based methods. We validate PixelSHAP in autonomous driving, highlighting its ability to enhance interpretability. Key challenges include segmentation sensitivity and object occlusion. Our open-source implementation facilitates further research.
title Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2503.06670