Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Shiming, Duan, Bowen, Khan, Salman, Khan, Fahad Shahbaz
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912457810771968
author Chen, Shiming
Duan, Bowen
Khan, Salman
Khan, Fahad Shahbaz
author_facet Chen, Shiming
Duan, Bowen
Khan, Salman
Khan, Fahad Shahbaz
contents Large-scale vision-language models (VLMs), such as CLIP, have achieved remarkable success in zero-shot learning (ZSL) by leveraging large-scale visual-text pair datasets. However, these methods often lack interpretability, as they compute the similarity between an entire query image and the embedded category words, making it difficult to explain their predictions. One approach to address this issue is to develop interpretable models by integrating language, where classifiers are built using discrete attributes, similar to human perception. This introduces a new challenge: how to effectively align local visual features with corresponding attributes based on pre-trained VLMs. To tackle this, we propose LaZSL, a locally-aligned vision-language model for interpretable ZSL. LaZSL employs local visual-semantic alignment via optimal transport to perform interaction between visual regions and their associated attributes, facilitating effective alignment and providing interpretable similarity without the need for additional training. Extensive experiments demonstrate that our method offers several advantages, including enhanced interpretability, improved accuracy, and strong domain generalization. Codes available at: https://github.com/shiming-chen/LaZSL.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23822
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
Chen, Shiming
Duan, Bowen
Khan, Salman
Khan, Fahad Shahbaz
Computer Vision and Pattern Recognition
Large-scale vision-language models (VLMs), such as CLIP, have achieved remarkable success in zero-shot learning (ZSL) by leveraging large-scale visual-text pair datasets. However, these methods often lack interpretability, as they compute the similarity between an entire query image and the embedded category words, making it difficult to explain their predictions. One approach to address this issue is to develop interpretable models by integrating language, where classifiers are built using discrete attributes, similar to human perception. This introduces a new challenge: how to effectively align local visual features with corresponding attributes based on pre-trained VLMs. To tackle this, we propose LaZSL, a locally-aligned vision-language model for interpretable ZSL. LaZSL employs local visual-semantic alignment via optimal transport to perform interaction between visual regions and their associated attributes, facilitating effective alignment and providing interpretable similarity without the need for additional training. Extensive experiments demonstrate that our method offers several advantages, including enhanced interpretability, improved accuracy, and strong domain generalization. Codes available at: https://github.com/shiming-chen/LaZSL.
title Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.23822