Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
Fuente:
arXiv
Saved in:
| Main Authors: | Hufe, Lorenz, Venhoff, Constantin, Purelku, Erblina, Dreyer, Maximilian, Lapuschkin, Sebastian, Samek, Wojciech |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
by: Dreyer, Maximilian, et al.
Published: (2024)
by: Dreyer, Maximilian, et al.
Published: (2024)
SCAM: A Real-World Typographic Robustness Evaluation for Multimodal Foundation Models
by: Westerhoff, Justus, et al.
Published: (2025)
by: Westerhoff, Justus, et al.
Published: (2025)
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
by: Dreyer, Maximilian, et al.
Published: (2023)
by: Dreyer, Maximilian, et al.
Published: (2023)
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
by: Bareeva, Dilyara, et al.
Published: (2024)
by: Bareeva, Dilyara, et al.
Published: (2024)
From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance
by: Dreyer, Maximilian, et al.
Published: (2025)
by: Dreyer, Maximilian, et al.
Published: (2025)
Pruning By Explaining Revisited: Optimizing Attribution Methods to Prune CNNs and Transformers
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2024)
by: Hatefi, Sayed Mohammad Vakilzadeh, et al.
Published: (2024)
Steering CLIP's vision transformer with sparse autoencoders
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
by: Erogullari, Eren, et al.
Published: (2025)
by: Erogullari, Eren, et al.
Published: (2025)
Ensuring Medical AI Safety: Interpretability-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data
by: Pahde, Frederik, et al.
Published: (2025)
by: Pahde, Frederik, et al.
Published: (2025)
Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
by: Pahde, Frederik, et al.
Published: (2022)
by: Pahde, Frederik, et al.
Published: (2022)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
by: Achtibat, Reduan, et al.
Published: (2024)
by: Achtibat, Reduan, et al.
Published: (2024)
Human-Centered Evaluation of XAI Methods
by: Dawoud, Karam, et al.
Published: (2023)
by: Dawoud, Karam, et al.
Published: (2023)
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples
by: Bouanani, Oussama, et al.
Published: (2026)
by: Bouanani, Oussama, et al.
Published: (2026)
Explainable concept mappings of MRI: Revealing the mechanisms underlying deep learning-based brain disease classification
by: Tinauer, Christian, et al.
Published: (2024)
by: Tinauer, Christian, et al.
Published: (2024)
XAI-guided Insulator Anomaly Detection for Imbalanced Datasets
by: Hoefler, Maximilian Andreas, et al.
Published: (2024)
by: Hoefler, Maximilian Andreas, et al.
Published: (2024)
Deep Learning-based Multi Project InP Wafer Simulation for Unsupervised Surface Defect Detection
by: Cantú, Emílio Dolgener, et al.
Published: (2025)
by: Cantú, Emílio Dolgener, et al.
Published: (2025)
Filtered-ViT: A Robust Defense Against Multiple Adversarial Patch Attacks
by: Khanal, Aja, et al.
Published: (2025)
by: Khanal, Aja, et al.
Published: (2025)
Concept-Based Masking: A Patch-Agnostic Defense Against Adversarial Patch Attacks
by: Mehrotra, Ayushi, et al.
Published: (2025)
by: Mehrotra, Ayushi, et al.
Published: (2025)
Robust Defense Strategies for Multimodal Contrastive Learning: Efficient Fine-tuning Against Backdoor Attacks
by: Hossain, Md. Iqbal, et al.
Published: (2025)
by: Hossain, Md. Iqbal, et al.
Published: (2025)
Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning
by: Qin, Chuan, et al.
Published: (2026)
by: Qin, Chuan, et al.
Published: (2026)
Synthetic Generation of Dermatoscopic Images with GAN and Closed-Form Factorization
by: Mekala, Rohan Reddy, et al.
Published: (2024)
by: Mekala, Rohan Reddy, et al.
Published: (2024)
A Real-Time Defense Against Object Vanishing Adversarial Patch Attacks for Object Detection in Autonomous Vehicles
by: Mu, Jaden
Published: (2024)
by: Mu, Jaden
Published: (2024)
Robust Vision-Language Models via Tensor Decomposition: A Defense Against Adversarial Attacks
by: Patel, Het, et al.
Published: (2025)
by: Patel, Het, et al.
Published: (2025)
CleanerCLIP: Fine-grained Counterfactual Semantic Augmentation for Backdoor Defense in Contrastive Learning
by: Xun, Yuan, et al.
Published: (2024)
by: Xun, Yuan, et al.
Published: (2024)
Defense-to-Attack: Bypassing Weak Defenses Enables Stronger Jailbreaks in Vision-Language Models
by: Zhao, Yunhan, et al.
Published: (2025)
by: Zhao, Yunhan, et al.
Published: (2025)
Test-Time Defense Against Adversarial Attacks via Stochastic Resonance of Latent Ensembles
by: Lao, Dong, et al.
Published: (2025)
by: Lao, Dong, et al.
Published: (2025)
Playing the network backward: A Game Theoretic Attribution Framework
by: Zimmermann, Jakob Paul, et al.
Published: (2026)
by: Zimmermann, Jakob Paul, et al.
Published: (2026)
Beyond Scalars: Concept-Based Alignment Analysis in Vision Transformers
by: Vielhaben, Johanna, et al.
Published: (2024)
by: Vielhaben, Johanna, et al.
Published: (2024)
Mechanistic understanding and validation of large AI models with SemanticLens
by: Dreyer, Maximilian, et al.
Published: (2025)
by: Dreyer, Maximilian, et al.
Published: (2025)
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
by: Cao, Yue, et al.
Published: (2024)
by: Cao, Yue, et al.
Published: (2024)
Defense That Attacks: How Robust Models Become Better Attackers
by: Awad, Mohamed, et al.
Published: (2025)
by: Awad, Mohamed, et al.
Published: (2025)
Certified Circuits: Stability Guarantees for Mechanistic Circuits
by: Anani, Alaa, et al.
Published: (2026)
by: Anani, Alaa, et al.
Published: (2026)
A Fresh Look at Sanity Checks for Saliency Maps
by: Hedström, Anna, et al.
Published: (2024)
by: Hedström, Anna, et al.
Published: (2024)
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
by: Cao, Anh-Quan, et al.
Published: (2024)
by: Cao, Anh-Quan, et al.
Published: (2024)
ECQ$^{\text{x}}$: Explainability-Driven Quantization for Low-Bit and Sparse DNNs
by: Becking, Daniel, et al.
Published: (2021)
by: Becking, Daniel, et al.
Published: (2021)
Model Inversion Attack Against Deep Hashing
by: Zhao, Dongdong, et al.
Published: (2025)
by: Zhao, Dongdong, et al.
Published: (2025)
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
by: Wu, Sihao, et al.
Published: (2025)
by: Wu, Sihao, et al.
Published: (2025)
Benchmarking Gaslighting Negation Attacks Against Reasoning Models
by: Zhu, Bin, et al.
Published: (2025)
by: Zhu, Bin, et al.
Published: (2025)
Attention-Based Real-Time Defenses for Physical Adversarial Attacks in Vision Applications
by: Rossolini, Giulio, et al.
Published: (2023)
by: Rossolini, Giulio, et al.
Published: (2023)
Similar Items
-
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
by: Dreyer, Maximilian, et al.
Published: (2024) -
SCAM: A Real-World Typographic Robustness Evaluation for Multimodal Foundation Models
by: Westerhoff, Justus, et al.
Published: (2025) -
Understanding the (Extra-)Ordinary: Validating Deep Model Decisions with Prototypical Concept-based Explanations
by: Dreyer, Maximilian, et al.
Published: (2023) -
Reactive Model Correction: Mitigating Harm to Task-Relevant Features via Conditional Bias Suppression
by: Bareeva, Dilyara, et al.
Published: (2024) -
From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance
by: Dreyer, Maximilian, et al.
Published: (2025)