Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hufe, Lorenz, Venhoff, Constantin, Purelku, Erblina, Dreyer, Maximilian, Lapuschkin, Sebastian, Samek, Wojciech
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915819373461504
author Hufe, Lorenz
Venhoff, Constantin
Purelku, Erblina
Dreyer, Maximilian
Lapuschkin, Sebastian
Samek, Wojciech
author_facet Hufe, Lorenz
Venhoff, Constantin
Purelku, Erblina
Dreyer, Maximilian
Lapuschkin, Sebastian
Samek, Wojciech
contents Typographic attacks exploit multi-modal systems by injecting text into images, leading to targeted misclassifications, malicious content generation and even Vision-Language Model jailbreaks. In this work, we analyze how CLIP vision encoders behave under typographic attacks, locating specialized attention heads in the latter half of the model's layers that causally extract and transmit typographic information to the cls token. Building on these insights, we introduce Dyslexify - a method to defend CLIP models against typographic attacks by selectively ablating a typographic circuit, consisting of attention heads. Without requiring finetuning, dyslexify improves performance by up to 22.06% on a typographic variant of ImageNet-100, while reducing standard ImageNet-100 accuracy by less than 1%, and demonstrate its utility in a medical foundation model for skin lesion diagnosis. Notably, our training-free approach remains competitive with current state-of-the-art typographic defenses that rely on finetuning. To this end, we release a family of dyslexic CLIP models which are significantly more robust against typographic attacks. These models serve as suitable drop-in replacements for a broad range of safety-critical applications, where the risks of text-based manipulation outweigh the utility of text recognition.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20570
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
Hufe, Lorenz
Venhoff, Constantin
Purelku, Erblina
Dreyer, Maximilian
Lapuschkin, Sebastian
Samek, Wojciech
Computer Vision and Pattern Recognition
Artificial Intelligence
Typographic attacks exploit multi-modal systems by injecting text into images, leading to targeted misclassifications, malicious content generation and even Vision-Language Model jailbreaks. In this work, we analyze how CLIP vision encoders behave under typographic attacks, locating specialized attention heads in the latter half of the model's layers that causally extract and transmit typographic information to the cls token. Building on these insights, we introduce Dyslexify - a method to defend CLIP models against typographic attacks by selectively ablating a typographic circuit, consisting of attention heads. Without requiring finetuning, dyslexify improves performance by up to 22.06% on a typographic variant of ImageNet-100, while reducing standard ImageNet-100 accuracy by less than 1%, and demonstrate its utility in a medical foundation model for skin lesion diagnosis. Notably, our training-free approach remains competitive with current state-of-the-art typographic defenses that rely on finetuning. To this end, we release a family of dyslexic CLIP models which are significantly more robust against typographic attacks. These models serve as suitable drop-in replacements for a broad range of safety-critical applications, where the risks of text-based manipulation outweigh the utility of text recognition.
title Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2508.20570