Embedding Hidden Adversarial Capabilities in Pre-Trained Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Beerens, Lucas, Higham, Desmond J.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913789262168064
author Beerens, Lucas
Higham, Desmond J.
author_facet Beerens, Lucas
Higham, Desmond J.
contents We introduce a new attack paradigm that embeds hidden adversarial capabilities directly into diffusion models via fine-tuning, without altering their observable behavior or requiring modifications during inference. Unlike prior approaches that target specific images or adjust the generation process to produce adversarial outputs, our method integrates adversarial functionality into the model itself. The resulting tampered model generates high-quality images indistinguishable from those of the original, yet these images cause misclassification in downstream classifiers at a high rate. The misclassification can be targeted to specific output classes. Users can employ this compromised model unaware of its embedded adversarial nature, as it functions identically to a standard diffusion model. We demonstrate the effectiveness and stealthiness of our approach, uncovering a covert attack vector that raises new security concerns. These findings expose a risk arising from the use of externally-supplied models and highlight the urgent need for robust model verification and defense mechanisms against hidden threats in generative models. The code is available at https://github.com/LucasBeerens/CRAFTed-Diffusion .
format Preprint
id arxiv_https___arxiv_org_abs_2504_08782
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Embedding Hidden Adversarial Capabilities in Pre-Trained Diffusion Models
Beerens, Lucas
Higham, Desmond J.
Machine Learning
Artificial Intelligence
Cryptography and Security
We introduce a new attack paradigm that embeds hidden adversarial capabilities directly into diffusion models via fine-tuning, without altering their observable behavior or requiring modifications during inference. Unlike prior approaches that target specific images or adjust the generation process to produce adversarial outputs, our method integrates adversarial functionality into the model itself. The resulting tampered model generates high-quality images indistinguishable from those of the original, yet these images cause misclassification in downstream classifiers at a high rate. The misclassification can be targeted to specific output classes. Users can employ this compromised model unaware of its embedded adversarial nature, as it functions identically to a standard diffusion model. We demonstrate the effectiveness and stealthiness of our approach, uncovering a covert attack vector that raises new security concerns. These findings expose a risk arising from the use of externally-supplied models and highlight the urgent need for robust model verification and defense mechanisms against hidden threats in generative models. The code is available at https://github.com/LucasBeerens/CRAFTed-Diffusion .
title Embedding Hidden Adversarial Capabilities in Pre-Trained Diffusion Models
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2504.08782