Guardado en:
Detalles Bibliográficos
Autores principales: Li, Zhiwei, Xue, Jiacheng, Wang, Weining, Liu, Ajian, Gao, Xingyu, Sun, Zhenan, Li, Qi
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:https://arxiv.org/abs/2605.15584
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917499135590400
author Li, Zhiwei
Xue, Jiacheng
Wang, Weining
Liu, Ajian
Gao, Xingyu
Sun, Zhenan
Li, Qi
author_facet Li, Zhiwei
Xue, Jiacheng
Wang, Weining
Liu, Ajian
Gao, Xingyu
Sun, Zhenan
Li, Qi
contents Vision-language models like CLIP have demonstrated remarkable zero-shot transfer capabilities. However, their susceptibility to imperceptible adversarial perturbations remains a critical security concern. While test-time defenses offer a pragmatic solution for deployed models, existing approaches typically rely on gradient-based optimization during inference, incurring significant computational overhead. In this paper, we revisit the role of data augmentation in CLIP robustness and observe that augmentations are not equally effective: specific augmentations consistently provide robust geometric cues that align with correct class semantics in the hyperspherical feature space. Based on this, we propose Adaptive Geodesic Correction (AGC), a training-free defense mechanism that requires no parameter updates. AGC identifies a reliable augmentation as a geometric anchor and corrects the input feature towards it, utilizing an adaptive step size to balance robustness against clean accuracy preservation. AGC achieves superior performance across eight fine-grained datasets and three CLIP backbones, improving average robust accuracy by 44.4\% over state-of-the-art baseline while delivering a 10$\times$ reduction in inference latency. Our findings reveal a fundamental geometric property of CLIP features, offering a highly efficient and effective paradigm for robust multimodal deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15584
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AGC: Adaptive Geodesic Correction for Adversarial Robustness on Vision-Language Models
Li, Zhiwei
Xue, Jiacheng
Wang, Weining
Liu, Ajian
Gao, Xingyu
Sun, Zhenan
Li, Qi
Computer Vision and Pattern Recognition
Vision-language models like CLIP have demonstrated remarkable zero-shot transfer capabilities. However, their susceptibility to imperceptible adversarial perturbations remains a critical security concern. While test-time defenses offer a pragmatic solution for deployed models, existing approaches typically rely on gradient-based optimization during inference, incurring significant computational overhead. In this paper, we revisit the role of data augmentation in CLIP robustness and observe that augmentations are not equally effective: specific augmentations consistently provide robust geometric cues that align with correct class semantics in the hyperspherical feature space. Based on this, we propose Adaptive Geodesic Correction (AGC), a training-free defense mechanism that requires no parameter updates. AGC identifies a reliable augmentation as a geometric anchor and corrects the input feature towards it, utilizing an adaptive step size to balance robustness against clean accuracy preservation. AGC achieves superior performance across eight fine-grained datasets and three CLIP backbones, improving average robust accuracy by 44.4\% over state-of-the-art baseline while delivering a 10$\times$ reduction in inference latency. Our findings reveal a fundamental geometric property of CLIP features, offering a highly efficient and effective paradigm for robust multimodal deployment.
title AGC: Adaptive Geodesic Correction for Adversarial Robustness on Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.15584