Diversifying Counterattacks: Orthogonal Exploration for Robust CLIP Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Chengze, Dong, Minjing, Shi, Xinli, Gui, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Survey of Adversarial Robustness in Multimodal Large Language Models
by: Jiang, Chengze, et al.
Published: (2025)
by: Jiang, Chengze, et al.
Published: (2025)
Improving Fast Adversarial Training Paradigm: An Example Taxonomy Perspective
by: Gui, Jie, et al.
Published: (2024)
by: Gui, Jie, et al.
Published: (2024)
Improving Fast Adversarial Training via Self-Knowledge Guidance
by: Jiang, Chengze, et al.
Published: (2024)
by: Jiang, Chengze, et al.
Published: (2024)
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
by: Xing, Songlong, et al.
Published: (2025)
by: Xing, Songlong, et al.
Published: (2025)
Revisiting Adversarial Training under Hyperspectral Image
by: Zhang, Weihua, et al.
Published: (2025)
by: Zhang, Weihua, et al.
Published: (2025)
Efficient Image-to-Image Diffusion Classifier for Adversarial Robustness
by: Mei, Hefei, et al.
Published: (2024)
by: Mei, Hefei, et al.
Published: (2024)
APC: Transferable and Efficient Adversarial Point Counterattack for Robust 3D Point Cloud Recognition
by: Jung, Geunyoung, et al.
Published: (2026)
by: Jung, Geunyoung, et al.
Published: (2026)
Learn to Preserve and Diversify: Parameter-Efficient Group with Orthogonal Regularization for Domain Generalization
by: Hu, Jiajun, et al.
Published: (2024)
by: Hu, Jiajun, et al.
Published: (2024)
Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation
by: Shi, Yuheng, et al.
Published: (2024)
by: Shi, Yuheng, et al.
Published: (2024)
Multi-Scale VMamba: Hierarchy in Hierarchy Visual State Space Model
by: Shi, Yuheng, et al.
Published: (2024)
by: Shi, Yuheng, et al.
Published: (2024)
Backdooring Self-Supervised Contrastive Learning by Noisy Alignment
by: Chen, Tuo, et al.
Published: (2025)
by: Chen, Tuo, et al.
Published: (2025)
Efficient Diffusion-Based 3D Human Pose Estimation with Hierarchical Temporal Pruning
by: Bi, Yuquan, et al.
Published: (2025)
by: Bi, Yuquan, et al.
Published: (2025)
ColorVein: Colorful Cancelable Vein Biometrics
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
by: Zhang, Jihai, et al.
Published: (2024)
by: Zhang, Jihai, et al.
Published: (2024)
Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception
by: Shi, Yuheng, et al.
Published: (2025)
by: Shi, Yuheng, et al.
Published: (2025)
VSSD: Vision Mamba with Non-Causal State Space Duality
by: Shi, Yuheng, et al.
Published: (2024)
by: Shi, Yuheng, et al.
Published: (2024)
Enhancing Robustness of Vision-Language Models through Orthogonality Learning and Self-Regularization
by: Li, Jinlong, et al.
Published: (2024)
by: Li, Jinlong, et al.
Published: (2024)
A Survey on Small Sample Imbalance Problem: Metrics, Feature Analysis, and Solutions
by: Zhao, Shuxian, et al.
Published: (2025)
by: Zhao, Shuxian, et al.
Published: (2025)
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference
by: Lan, Mengcheng, et al.
Published: (2024)
by: Lan, Mengcheng, et al.
Published: (2024)
Feature Clipping for Uncertainty Calibration
by: Tao, Linwei, et al.
Published: (2024)
by: Tao, Linwei, et al.
Published: (2024)
NEARL-CLIP: Interacted Query Adaptation with Orthogonal Regularization for Medical Vision-Language Understanding
by: Peng, Zelin, et al.
Published: (2025)
by: Peng, Zelin, et al.
Published: (2025)
MEDiC: Multi-objective Exploration of Distillation from CLIP
by: Georgiou, Konstantinos, et al.
Published: (2026)
by: Georgiou, Konstantinos, et al.
Published: (2026)
Exploring the Coordination of Frequency and Attention in Masked Image Modeling
by: Gui, Jie, et al.
Published: (2022)
by: Gui, Jie, et al.
Published: (2022)
CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian Splatting
by: Jiao, Siyu, et al.
Published: (2024)
by: Jiao, Siyu, et al.
Published: (2024)
Long-CLIP: Unlocking the Long-Text Capability of CLIP
by: Zhang, Beichen, et al.
Published: (2024)
by: Zhang, Beichen, et al.
Published: (2024)
Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models
by: Shi, Yuheng, et al.
Published: (2026)
by: Shi, Yuheng, et al.
Published: (2026)
Diversifying Query: Region-Guided Transformer for Temporal Sentence Grounding
by: Sun, Xiaolong, et al.
Published: (2024)
by: Sun, Xiaolong, et al.
Published: (2024)
VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language Models
by: Mei, Hefei, et al.
Published: (2025)
by: Mei, Hefei, et al.
Published: (2025)
PA-Attack: Guiding Gray-Box Attacks on LVLM Vision Encoders with Prototypes and Attention
by: Mei, Hefei, et al.
Published: (2026)
by: Mei, Hefei, et al.
Published: (2026)
Contrastive Spectral Rectification: Test-Time Defense towards Zero-shot Adversarial Robustness of CLIP
by: Nie, Sen, et al.
Published: (2026)
by: Nie, Sen, et al.
Published: (2026)
CLIP-AGIQA: Boosting the Performance of AI-Generated Image Quality Assessment with CLIP
by: Tang, Zhenchen, et al.
Published: (2024)
by: Tang, Zhenchen, et al.
Published: (2024)
Multimodal CLIP Inference for Meta-Few-Shot Image Classification
by: Ferragu, Constance, et al.
Published: (2024)
by: Ferragu, Constance, et al.
Published: (2024)
CLIP-FSAC++: Few-Shot Anomaly Classification with Anomaly Descriptor Based on CLIP
by: Zuo, Zuo, et al.
Published: (2024)
by: Zuo, Zuo, et al.
Published: (2024)
Toward a Holistic Evaluation of Robustness in CLIP Models
by: Tu, Weijie, et al.
Published: (2024)
by: Tu, Weijie, et al.
Published: (2024)
Control-CLIP: Decoupling Category and Style Guidance in CLIP for Specific-Domain Generation
by: Jia, Zexi, et al.
Published: (2025)
by: Jia, Zexi, et al.
Published: (2025)
Quantified Task Misalignment to Inform PEFT: An Exploration of Domain Generalization and Catastrophic Forgetting in CLIP
by: Niss, Laura, et al.
Published: (2024)
by: Niss, Laura, et al.
Published: (2024)
Mitigating Object Hallucinations in Large Vision-Language Models via Attention Calibration
by: Zhu, Younan, et al.
Published: (2025)
by: Zhu, Younan, et al.
Published: (2025)
Beyond One-Hot Labels: Semantic Mixing for Model Calibration
by: Luo, Haoyang, et al.
Published: (2025)
by: Luo, Haoyang, et al.
Published: (2025)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
by: Yang, Kaicheng, et al.
Published: (2024)
by: Yang, Kaicheng, et al.
Published: (2024)
Robust Orthogonal NMF with Label Propagation for Image Clustering
by: Liu, Jingjing, et al.
Published: (2025)
by: Liu, Jingjing, et al.
Published: (2025)
Similar Items
-
Survey of Adversarial Robustness in Multimodal Large Language Models
by: Jiang, Chengze, et al.
Published: (2025) -
Improving Fast Adversarial Training Paradigm: An Example Taxonomy Perspective
by: Gui, Jie, et al.
Published: (2024) -
Improving Fast Adversarial Training via Self-Knowledge Guidance
by: Jiang, Chengze, et al.
Published: (2024) -
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
by: Xing, Songlong, et al.
Published: (2025) -
Revisiting Adversarial Training under Hyperspectral Image
by: Zhang, Weihua, et al.
Published: (2025)