Robustness in Both Domains: CLIP Needs a Robust Text Encoder
Fuente:
arXiv
Salvato in:
| Autori principali: | Rocamora, Elias Abad, Schlarmann, Christian, Singh, Naman Deep, Wu, Yongtao, Hein, Matthias, Cevher, Volkan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models
di: Schlarmann, Christian, et al.
Pubblicazione: (2024)
di: Schlarmann, Christian, et al.
Pubblicazione: (2024)
Adversarially Robust CLIP Models Can Induce Better (Robust) Perceptual Metrics
di: Croce, Francesco, et al.
Pubblicazione: (2025)
di: Croce, Francesco, et al.
Pubblicazione: (2025)
Membership Inference Attacks against Large Vision-Language Models
di: Li, Zhan, et al.
Pubblicazione: (2024)
di: Li, Zhan, et al.
Pubblicazione: (2024)
Advancing Compositional Awareness in CLIP with Efficient Fine-Tuning
di: Peleg, Amit, et al.
Pubblicazione: (2025)
di: Peleg, Amit, et al.
Pubblicazione: (2025)
Perturb and Recover: Fine-tuning for Effective Backdoor Removal from CLIP
di: Singh, Naman Deep, et al.
Pubblicazione: (2024)
di: Singh, Naman Deep, et al.
Pubblicazione: (2024)
Visual Memory Injection Attacks for Multi-Turn Conversations
di: Schlarmann, Christian, et al.
Pubblicazione: (2026)
di: Schlarmann, Christian, et al.
Pubblicazione: (2026)
Certified Robustness Under Bounded Levenshtein Distance
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2025)
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2025)
Towards Reliable Evaluation and Fast Training of Robust Semantic Segmentation Models
di: Croce, Francesco, et al.
Pubblicazione: (2023)
di: Croce, Francesco, et al.
Pubblicazione: (2023)
DiffCAP: Diffusion-based Cumulative Adversarial Purification for Vision Language Models
di: Fu, Jia, et al.
Pubblicazione: (2025)
di: Fu, Jia, et al.
Pubblicazione: (2025)
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
di: Schlarmann, Christian, et al.
Pubblicazione: (2025)
di: Schlarmann, Christian, et al.
Pubblicazione: (2025)
Revisiting Character-level Adversarial Attacks for Language Models
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2024)
di: Rocamora, Elias Abad, et al.
Pubblicazione: (2024)
Fine-tuning CLIP Text Encoders with Two-step Paraphrasing
di: Kim, Hyunjae, et al.
Pubblicazione: (2024)
di: Kim, Hyunjae, et al.
Pubblicazione: (2024)
Occlusion Robustness of CLIP for Military Vehicle Classification
di: van Woerden, Jan Erik, et al.
Pubblicazione: (2025)
di: van Woerden, Jan Erik, et al.
Pubblicazione: (2025)
Single-pass Detection of Jailbreaking Input in Large Language Models
di: Candogan, Leyla Naz, et al.
Pubblicazione: (2025)
di: Candogan, Leyla Naz, et al.
Pubblicazione: (2025)
Going beyond Compositions, DDPMs Can Produce Zero-Shot Interpolations
di: Deschenaux, Justin, et al.
Pubblicazione: (2024)
di: Deschenaux, Justin, et al.
Pubblicazione: (2024)
Mind the Detail: Uncovering Clinically Relevant Image Details in Accelerated MRI with Semantically Diverse Reconstructions
di: Morshuis, Jan Nikolas, et al.
Pubblicazione: (2025)
di: Morshuis, Jan Nikolas, et al.
Pubblicazione: (2025)
Expanding Event Modality Applications through a Robust CLIP-Based Encoder
di: Jeong, Sungheon, et al.
Pubblicazione: (2024)
di: Jeong, Sungheon, et al.
Pubblicazione: (2024)
Adversarial Training on Purification (AToP): Advancing Both Robustness and Generalization
di: Lin, Guang, et al.
Pubblicazione: (2024)
di: Lin, Guang, et al.
Pubblicazione: (2024)
VajraV1 -- The most accurate Real Time Object Detector of the YOLO family
di: Makkar, Naman Balbir Singh
Pubblicazione: (2025)
di: Makkar, Naman Balbir Singh
Pubblicazione: (2025)
Helping CLIP See Both the Forest and the Trees: A Decomposition and Description Approach
di: Xue, Leyan, et al.
Pubblicazione: (2025)
di: Xue, Leyan, et al.
Pubblicazione: (2025)
Noise-Robust AV-ASR Using Visual Features Both in the Whisper Encoder and Decoder
di: Li, Zhengyang, et al.
Pubblicazione: (2026)
di: Li, Zhengyang, et al.
Pubblicazione: (2026)
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
di: Yu, Lu, et al.
Pubblicazione: (2024)
di: Yu, Lu, et al.
Pubblicazione: (2024)
Investigating the Semantic Robustness of CLIP-based Zero-Shot Anomaly Segmentation
di: Stangl, Kevin, et al.
Pubblicazione: (2024)
di: Stangl, Kevin, et al.
Pubblicazione: (2024)
The Last Mile to Supervised Performance: Semi-Supervised Domain Adaptation for Semantic Segmentation
di: Morales-Brotons, Daniel, et al.
Pubblicazione: (2024)
di: Morales-Brotons, Daniel, et al.
Pubblicazione: (2024)
DesCLIP: Robust Continual Learning via General Attribute Descriptions for VLM-Based Visual Recognition
di: He, Chiyuan, et al.
Pubblicazione: (2025)
di: He, Chiyuan, et al.
Pubblicazione: (2025)
CLIPure: Purification in Latent Space via CLIP for Adversarially Robust Zero-Shot Classification
di: Zhang, Mingkun, et al.
Pubblicazione: (2025)
di: Zhang, Mingkun, et al.
Pubblicazione: (2025)
Certified Zeroth-order Black-Box Defense with Robust UNet Denoiser
di: Verma, Astha, et al.
Pubblicazione: (2023)
di: Verma, Astha, et al.
Pubblicazione: (2023)
Interpreting CLIP: Insights on the Robustness to ImageNet Distribution Shifts
di: Crabbé, Jonathan, et al.
Pubblicazione: (2023)
di: Crabbé, Jonathan, et al.
Pubblicazione: (2023)
Securing Vision-Language Models with a Robust Encoder Against Jailbreak and Adversarial Attacks
di: Hossain, Md Zarif, et al.
Pubblicazione: (2024)
di: Hossain, Md Zarif, et al.
Pubblicazione: (2024)
Parrot Captions Teach CLIP to Spot Text
di: Lin, Yiqi, et al.
Pubblicazione: (2023)
di: Lin, Yiqi, et al.
Pubblicazione: (2023)
Prompt Group-Aware Training for Robust Text-Guided Nuclei Segmentation
di: Wu, Yonghuang, et al.
Pubblicazione: (2026)
di: Wu, Yonghuang, et al.
Pubblicazione: (2026)
Zoom-shot: Fast and Efficient Unsupervised Zero-Shot Transfer of CLIP to Vision Encoders with Multimodal Loss
di: Shipard, Jordan, et al.
Pubblicazione: (2024)
di: Shipard, Jordan, et al.
Pubblicazione: (2024)
Locating Demographic Bias at the Attention-Head Level in CLIP's Vision Encoder
di: Yasser, Alaa, et al.
Pubblicazione: (2026)
di: Yasser, Alaa, et al.
Pubblicazione: (2026)
Designing a Robust Radiology Report Generation System
di: Singh, Sonit
Pubblicazione: (2024)
di: Singh, Sonit
Pubblicazione: (2024)
Robust Domain Generalization for Multi-modal Object Recognition
di: Qiao, Yuxin, et al.
Pubblicazione: (2024)
di: Qiao, Yuxin, et al.
Pubblicazione: (2024)
CLIP is All You Need for Human-like Semantic Representations in Stable Diffusion
di: Braunstein, Cameron, et al.
Pubblicazione: (2025)
di: Braunstein, Cameron, et al.
Pubblicazione: (2025)
CT-CLIP: A Multi-modal Fusion Framework for Robust Apple Leaf Disease Recognition in Complex Environments
di: Liu, Lemin, et al.
Pubblicazione: (2025)
di: Liu, Lemin, et al.
Pubblicazione: (2025)
Is Your Text-to-Image Model Robust to Caption Noise?
di: Yu, Weichen, et al.
Pubblicazione: (2024)
di: Yu, Weichen, et al.
Pubblicazione: (2024)
Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
di: Che, Chang, et al.
Pubblicazione: (2024)
di: Che, Chang, et al.
Pubblicazione: (2024)
Interpreting CLIP's Image Representation via Text-Based Decomposition
di: Gandelsman, Yossi, et al.
Pubblicazione: (2023)
di: Gandelsman, Yossi, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models
di: Schlarmann, Christian, et al.
Pubblicazione: (2024) -
Adversarially Robust CLIP Models Can Induce Better (Robust) Perceptual Metrics
di: Croce, Francesco, et al.
Pubblicazione: (2025) -
Membership Inference Attacks against Large Vision-Language Models
di: Li, Zhan, et al.
Pubblicazione: (2024) -
Advancing Compositional Awareness in CLIP with Efficient Fine-Tuning
di: Peleg, Amit, et al.
Pubblicazione: (2025) -
Perturb and Recover: Fine-tuning for Effective Backdoor Removal from CLIP
di: Singh, Naman Deep, et al.
Pubblicazione: (2024)