Backdoor Directions in Vision Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Karayalcin, Sengim, Krcek, Marina, Chen, Pin-Yu, Picek, Stjepan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915854089715712
author Karayalcin, Sengim
Krcek, Marina
Chen, Pin-Yu
Picek, Stjepan
author_facet Karayalcin, Sengim
Krcek, Marina
Chen, Pin-Yu
Picek, Stjepan
contents This paper investigates how Backdoor Attacks are represented within Vision Transformers (ViTs). By assuming knowledge of the trigger, we identify a specific ``trigger direction'' in the model's activations that corresponds to the internal representation of the trigger. We confirm the causal role of this linear direction by showing that interventions in both activation and parameter space consistently modulate the model's backdoor behavior across multiple datasets and attack types. Using this direction as a diagnostic tool, we trace how backdoor features are processed across layers. Our analysis reveals distinct qualitative differences: static-patch triggers follow a different internal logic than stealthy, distributed triggers. We further examine the link between backdoors and adversarial attacks, specifically testing whether PGD-based perturbations (de-)activate the identified trigger mechanism. Finally, we propose a data-free, weight-based detection scheme for stealthy-trigger attacks. Our findings show that mechanistic interpretability offers a robust framework for diagnosing and addressing security vulnerabilities in computer vision.
format Preprint
id arxiv_https___arxiv_org_abs_2603_10806
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Backdoor Directions in Vision Transformers
Karayalcin, Sengim
Krcek, Marina
Chen, Pin-Yu
Picek, Stjepan
Computer Vision and Pattern Recognition
Cryptography and Security
This paper investigates how Backdoor Attacks are represented within Vision Transformers (ViTs). By assuming knowledge of the trigger, we identify a specific ``trigger direction'' in the model's activations that corresponds to the internal representation of the trigger. We confirm the causal role of this linear direction by showing that interventions in both activation and parameter space consistently modulate the model's backdoor behavior across multiple datasets and attack types. Using this direction as a diagnostic tool, we trace how backdoor features are processed across layers. Our analysis reveals distinct qualitative differences: static-patch triggers follow a different internal logic than stealthy, distributed triggers. We further examine the link between backdoors and adversarial attacks, specifically testing whether PGD-based perturbations (de-)activate the identified trigger mechanism. Finally, we propose a data-free, weight-based detection scheme for stealthy-trigger attacks. Our findings show that mechanistic interpretability offers a robust framework for diagnosing and addressing security vulnerabilities in computer vision.
title Backdoor Directions in Vision Transformers
topic Computer Vision and Pattern Recognition
Cryptography and Security
url https://arxiv.org/abs/2603.10806