Semantic-level Backdoor Attack against Text-to-Image Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Tianxin, Jiang, Wenbo, Chen, Hongqiao, Zheng, Zhirun, Huang, Cheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913165012369408
author Chen, Tianxin
Jiang, Wenbo
Chen, Hongqiao
Zheng, Zhirun
Huang, Cheng
author_facet Chen, Tianxin
Jiang, Wenbo
Chen, Hongqiao
Zheng, Zhirun
Huang, Cheng
contents Text-to-image (T2I) diffusion models are widely adopted for their strong generative capabilities, yet remain vulnerable to backdoor attacks. Existing attacks typically rely on fixed textual triggers and single-entity backdoor targets, making them highly susceptible to enumeration-based input defenses and attention-consistency detection. In this work, we propose Semantic-level Backdoor Attack (SemBD), which introduces representation-level triggers based on continuous semantic regions rather than discrete textual patterns. SemBD implants such semantic backdoors by distillation-based editing of the key and value projection matrices in cross-attention layers, enabling semantically equivalent but textually diverse prompts to activate the backdoor. To further enhance stealthiness, SemBD incorporates a semantic regularization to prevent unintended activation under incomplete semantics, as well as multi-entity backdoor targets that avoid highly consistent cross-attention patterns. Extensive experiments demonstrate that SemBD achieves a 100% attack success rate while maintaining strong robustness against state-of-the-art input-level defenses. Our code is available at https://github.com/DPAS-Lab/SemBD/.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04898
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Semantic-level Backdoor Attack against Text-to-Image Diffusion Models
Chen, Tianxin
Jiang, Wenbo
Chen, Hongqiao
Zheng, Zhirun
Huang, Cheng
Cryptography and Security
Artificial Intelligence
Text-to-image (T2I) diffusion models are widely adopted for their strong generative capabilities, yet remain vulnerable to backdoor attacks. Existing attacks typically rely on fixed textual triggers and single-entity backdoor targets, making them highly susceptible to enumeration-based input defenses and attention-consistency detection. In this work, we propose Semantic-level Backdoor Attack (SemBD), which introduces representation-level triggers based on continuous semantic regions rather than discrete textual patterns. SemBD implants such semantic backdoors by distillation-based editing of the key and value projection matrices in cross-attention layers, enabling semantically equivalent but textually diverse prompts to activate the backdoor. To further enhance stealthiness, SemBD incorporates a semantic regularization to prevent unintended activation under incomplete semantics, as well as multi-entity backdoor targets that avoid highly consistent cross-attention patterns. Extensive experiments demonstrate that SemBD achieves a 100% attack success rate while maintaining strong robustness against state-of-the-art input-level defenses. Our code is available at https://github.com/DPAS-Lab/SemBD/.
title Semantic-level Backdoor Attack against Text-to-Image Diffusion Models
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2602.04898