Zero-Shot Robustness of Vision Language Models Via Confidence-Aware Weighting

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Naghavian, Nikoo, Tavassolipour, Mostafa
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912626217320448
author Naghavian, Nikoo
Tavassolipour, Mostafa
author_facet Naghavian, Nikoo
Tavassolipour, Mostafa
contents Vision-language models like CLIP demonstrate impressive zero-shot generalization but remain highly vulnerable to adversarial attacks. In this work, we propose Confidence-Aware Weighting (CAW) to enhance zero-shot robustness in vision-language models. CAW consists of two components: (1) a Confidence-Aware loss that prioritizes uncertain adversarial examples by scaling the KL divergence between clean and adversarial predictions, and (2) a feature alignment regularization that preserves semantic consistency by minimizing the distance between frozen and fine-tuned image encoder features on adversarial inputs. These components work jointly to improve both clean and robust accuracy without sacrificing generalization. Extensive experiments on TinyImageNet and 14 additional datasets show that CAW outperforms recent methods such as PMG-AFT and TGA-ZSR under strong attacks like AutoAttack, while using less memory.
format Preprint
id arxiv_https___arxiv_org_abs_2510_02913
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Zero-Shot Robustness of Vision Language Models Via Confidence-Aware Weighting
Naghavian, Nikoo
Tavassolipour, Mostafa
Computer Vision and Pattern Recognition
Vision-language models like CLIP demonstrate impressive zero-shot generalization but remain highly vulnerable to adversarial attacks. In this work, we propose Confidence-Aware Weighting (CAW) to enhance zero-shot robustness in vision-language models. CAW consists of two components: (1) a Confidence-Aware loss that prioritizes uncertain adversarial examples by scaling the KL divergence between clean and adversarial predictions, and (2) a feature alignment regularization that preserves semantic consistency by minimizing the distance between frozen and fine-tuned image encoder features on adversarial inputs. These components work jointly to improve both clean and robust accuracy without sacrificing generalization. Extensive experiments on TinyImageNet and 14 additional datasets show that CAW outperforms recent methods such as PMG-AFT and TGA-ZSR under strong attacks like AutoAttack, while using less memory.
title Zero-Shot Robustness of Vision Language Models Via Confidence-Aware Weighting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.02913