Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Xing, Songlong, Wang, Weijie, Zhao, Zhengyu, Gu, Jindong, Torr, Philip, Sebe, Nicu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
by: Xing, Songlong, et al.
Published: (2025)
by: Xing, Songlong, et al.
Published: (2025)
PoInit-of-View: Poisoning Initialization of Views Transfers Across Multiple 3D Reconstruction Systems
by: Wang, Weijie, et al.
Published: (2026)
by: Wang, Weijie, et al.
Published: (2026)
An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models
by: Luo, Haochen, et al.
Published: (2024)
by: Luo, Haochen, et al.
Published: (2024)
Hyper Adversarial Tuning for Boosting Adversarial Robustness of Pretrained Large Vision Models
by: Lv, Kangtao, et al.
Published: (2024)
by: Lv, Kangtao, et al.
Published: (2024)
Improving Adversarial Transferability via Model Alignment
by: Ma, Avery, et al.
Published: (2023)
by: Ma, Avery, et al.
Published: (2023)
Hierarchically Robust Zero-shot Vision-language Models
by: Dong, Junhao, et al.
Published: (2026)
by: Dong, Junhao, et al.
Published: (2026)
Enhancing Robustness of Vision-Language Models through Orthogonality Learning and Self-Regularization
by: Li, Jinlong, et al.
Published: (2024)
by: Li, Jinlong, et al.
Published: (2024)
ZeroReg: Zero-Shot Point Cloud Registration with Foundation Models
by: Wang, Weijie, et al.
Published: (2023)
by: Wang, Weijie, et al.
Published: (2023)
Influencer Backdoor Attack on Semantic Segmentation
by: Lan, Haoheng, et al.
Published: (2023)
by: Lan, Haoheng, et al.
Published: (2023)
Rethinking the Learning Paradigm for Facial Expression Recognition
by: Wang, Weijie, et al.
Published: (2022)
by: Wang, Weijie, et al.
Published: (2022)
Causal Disentanglement for Robust Long-tail Medical Image Generation
by: Nie, Weizhi, et al.
Published: (2025)
by: Nie, Weizhi, et al.
Published: (2025)
FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized Segmentation
by: Zhao, Dong, et al.
Published: (2025)
by: Zhao, Dong, et al.
Published: (2025)
Which Model Generated This Image? A Model-Agnostic Approach for Origin Attribution
by: Liu, Fengyuan, et al.
Published: (2024)
by: Liu, Fengyuan, et al.
Published: (2024)
Learning Visual Prompts for Guiding the Attention of Vision Transformers
by: Rezaei, Razieh, et al.
Published: (2024)
by: Rezaei, Razieh, et al.
Published: (2024)
Vision+X: A Survey on Multimodal Learning in the Light of Data
by: Zhu, Ye, et al.
Published: (2022)
by: Zhu, Ye, et al.
Published: (2022)
Transferable-guided Attention Is All You Need for Video Domain Adaptation
by: Sacilotti, André, et al.
Published: (2024)
by: Sacilotti, André, et al.
Published: (2024)
Large Language Models for Multimodal Deformable Image Registration
by: Ma, Mingrui, et al.
Published: (2024)
by: Ma, Mingrui, et al.
Published: (2024)
BackdoorIDS: Zero-shot Backdoor Detection for Pretrained Vision Encoder
by: Huang, Siquan, et al.
Published: (2026)
by: Huang, Siquan, et al.
Published: (2026)
Reliable Few-shot Learning under Dual Noises
by: Zhang, Ji, et al.
Published: (2025)
by: Zhang, Ji, et al.
Published: (2025)
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
by: Wu, Yixuan, et al.
Published: (2025)
by: Wu, Yixuan, et al.
Published: (2025)
Self-Discovering Interpretable Diffusion Latent Directions for Responsible Text-to-Image Generation
by: Li, Hang, et al.
Published: (2023)
by: Li, Hang, et al.
Published: (2023)
As Firm As Their Foundations: Can open-sourced foundation models be used to create adversarial examples for downstream tasks?
by: Hu, Anjun, et al.
Published: (2024)
by: Hu, Anjun, et al.
Published: (2024)
Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
Generalizable Knowledge Distillation from Vision Foundation Models for Semantic Segmentation
by: Lv, Chonghua, et al.
Published: (2026)
by: Lv, Chonghua, et al.
Published: (2026)
Asymmetric GANs for Image-to-Image Translation
by: Tang, Hao, et al.
Published: (2019)
by: Tang, Hao, et al.
Published: (2019)
Boosting Audio-visual Zero-shot Learning with Large Language Models
by: Chen, Haoxing, et al.
Published: (2023)
by: Chen, Haoxing, et al.
Published: (2023)
FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors
by: Li, Chenxi, et al.
Published: (2025)
by: Li, Chenxi, et al.
Published: (2025)
Fully-Geometric Cross-Attention for Point Cloud Registration
by: Wang, Weijie, et al.
Published: (2025)
by: Wang, Weijie, et al.
Published: (2025)
Improving Adversarial Transferability in MLLMs via Dynamic Vision-Language Alignment Attack
by: Gu, Chenhe, et al.
Published: (2025)
by: Gu, Chenhe, et al.
Published: (2025)
A Closer Look at Conditional Prompt Tuning for Vision-Language Models
by: Zhang, Ji, et al.
Published: (2025)
by: Zhang, Ji, et al.
Published: (2025)
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
by: Li, Jinlong, et al.
Published: (2026)
by: Li, Jinlong, et al.
Published: (2026)
Reverse Personalization
by: Kung, Han-Wei, et al.
Published: (2025)
by: Kung, Han-Wei, et al.
Published: (2025)
RankFeat&RankWeight: Rank-1 Feature/Weight Removal for Out-of-distribution Detection
by: Song, Yue, et al.
Published: (2023)
by: Song, Yue, et al.
Published: (2023)
FullLoRA: Efficiently Boosting the Robustness of Pretrained Vision Transformers
by: Yuan, Zheng, et al.
Published: (2024)
by: Yuan, Zheng, et al.
Published: (2024)
MAKE: Multi-Aspect Knowledge-Enhanced Vision-Language Pretraining for Zero-shot Dermatological Assessment
by: Yan, Siyuan, et al.
Published: (2025)
by: Yan, Siyuan, et al.
Published: (2025)
Anchor-based Robust Finetuning of Vision-Language Models
by: Han, Jinwei, et al.
Published: (2024)
by: Han, Jinwei, et al.
Published: (2024)
Graph Transformer GANs with Graph Masked Modeling for Architectural Layout Generation
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
POCI-Diff: Position Objects Consistently and Interactively with 3D-Layout Guided Diffusion
by: Rigo, Andrea, et al.
Published: (2026)
by: Rigo, Andrea, et al.
Published: (2026)
Safe Vision-Language Models via Unsafe Weights Manipulation
by: D'Incà, Moreno, et al.
Published: (2025)
by: D'Incà, Moreno, et al.
Published: (2025)
UVMap-ID: A Controllable and Personalized UV Map Generative Model
by: Wang, Weijie, et al.
Published: (2024)
by: Wang, Weijie, et al.
Published: (2024)
Similar Items
-
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
by: Xing, Songlong, et al.
Published: (2025) -
PoInit-of-View: Poisoning Initialization of Views Transfers Across Multiple 3D Reconstruction Systems
by: Wang, Weijie, et al.
Published: (2026) -
An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models
by: Luo, Haochen, et al.
Published: (2024) -
Hyper Adversarial Tuning for Boosting Adversarial Robustness of Pretrained Large Vision Models
by: Lv, Kangtao, et al.
Published: (2024) -
Improving Adversarial Transferability via Model Alignment
by: Ma, Avery, et al.
Published: (2023)