DiSa: Directional Saliency-Aware Prompt Learning for Generalizable Vision-Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Talemi, Niloufar Alipour, Kashiani, Hossein, Nowdeh, Hossein R., Afghah, Fatemeh
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916758780116992
author Talemi, Niloufar Alipour
Kashiani, Hossein
Nowdeh, Hossein R.
Afghah, Fatemeh
author_facet Talemi, Niloufar Alipour
Kashiani, Hossein
Nowdeh, Hossein R.
Afghah, Fatemeh
contents Prompt learning has emerged as a powerful paradigm for adapting vision-language models such as CLIP to downstream tasks. However, existing methods often overfit to seen data, leading to significant performance degradation when generalizing to novel classes or unseen domains. To address this limitation, we propose DiSa, a Directional Saliency-Aware Prompt Learning framework that integrates two complementary regularization strategies to enhance generalization. First, our Cross-Interactive Regularization (CIR) fosters cross-modal alignment by enabling cooperative learning between prompted and frozen encoders. Within CIR, a saliency-aware masking strategy guides the image encoder to prioritize semantically critical image regions, reducing reliance on less informative patches. Second, we introduce a directional regularization strategy that aligns visual embeddings with class-wise prototype features in a directional manner to prioritize consistency in feature orientation over strict proximity. This approach ensures robust generalization by leveraging stable prototype directions derived from class-mean statistics. Extensive evaluations on 11 diverse image classification benchmarks demonstrate that DiSa consistently outperforms state-of-the-art prompt learning methods across various settings, including base-to-novel generalization, cross-dataset transfer, domain generalization, and few-shot learning.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19373
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiSa: Directional Saliency-Aware Prompt Learning for Generalizable Vision-Language Models
Talemi, Niloufar Alipour
Kashiani, Hossein
Nowdeh, Hossein R.
Afghah, Fatemeh
Computer Vision and Pattern Recognition
Prompt learning has emerged as a powerful paradigm for adapting vision-language models such as CLIP to downstream tasks. However, existing methods often overfit to seen data, leading to significant performance degradation when generalizing to novel classes or unseen domains. To address this limitation, we propose DiSa, a Directional Saliency-Aware Prompt Learning framework that integrates two complementary regularization strategies to enhance generalization. First, our Cross-Interactive Regularization (CIR) fosters cross-modal alignment by enabling cooperative learning between prompted and frozen encoders. Within CIR, a saliency-aware masking strategy guides the image encoder to prioritize semantically critical image regions, reducing reliance on less informative patches. Second, we introduce a directional regularization strategy that aligns visual embeddings with class-wise prototype features in a directional manner to prioritize consistency in feature orientation over strict proximity. This approach ensures robust generalization by leveraging stable prototype directions derived from class-mean statistics. Extensive evaluations on 11 diverse image classification benchmarks demonstrate that DiSa consistently outperforms state-of-the-art prompt learning methods across various settings, including base-to-novel generalization, cross-dataset transfer, domain generalization, and few-shot learning.
title DiSa: Directional Saliency-Aware Prompt Learning for Generalizable Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.19373