When Backdoors Go Beyond Triggers: Semantic Drift in Diffusion Models Under Encoder Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Shenyang, Zhu, Liuwan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912921337987072
author Chen, Shenyang
Zhu, Liuwan
author_facet Chen, Shenyang
Zhu, Liuwan
contents Standard evaluations of backdoor attacks on text-to-image (T2I) models primarily measure trigger activation and visual fidelity. We challenge this paradigm, demonstrating that encoder-side poisoning induces persistent, trigger-free semantic corruption that fundamentally reshapes the representation manifold. We trace this vulnerability to a geometric mechanism: a Jacobian-based analysis reveals that backdoors act as low-rank, target-centered deformations that amplify local sensitivity, causing distortion to propagate coherently across semantic neighborhoods. To rigorously quantify this structural degradation, we introduce SEMAD (Semantic Alignment and Drift), a diagnostic framework that measures both internal embedding drift and downstream functional misalignment. Our findings, validated across diffusion and contrastive paradigms, expose the deep structural risks of encoder poisoning and highlight the necessity of geometric audits beyond simple attack success rates.
format Preprint
id arxiv_https___arxiv_org_abs_2602_20193
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When Backdoors Go Beyond Triggers: Semantic Drift in Diffusion Models Under Encoder Attacks
Chen, Shenyang
Zhu, Liuwan
Cryptography and Security
Artificial Intelligence
Standard evaluations of backdoor attacks on text-to-image (T2I) models primarily measure trigger activation and visual fidelity. We challenge this paradigm, demonstrating that encoder-side poisoning induces persistent, trigger-free semantic corruption that fundamentally reshapes the representation manifold. We trace this vulnerability to a geometric mechanism: a Jacobian-based analysis reveals that backdoors act as low-rank, target-centered deformations that amplify local sensitivity, causing distortion to propagate coherently across semantic neighborhoods. To rigorously quantify this structural degradation, we introduce SEMAD (Semantic Alignment and Drift), a diagnostic framework that measures both internal embedding drift and downstream functional misalignment. Our findings, validated across diffusion and contrastive paradigms, expose the deep structural risks of encoder poisoning and highlight the necessity of geometric audits beyond simple attack success rates.
title When Backdoors Go Beyond Triggers: Semantic Drift in Diffusion Models Under Encoder Attacks
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2602.20193