Patronus: Safeguarding Text-to-Image Models against White-Box Adversaries

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Xinfeng, Pang, Shengyuan, Wu, Jialin, Deng, Jiangyi, Zhong, Huanlong, Chen, Yanjiao, Zhang, Jie, Xu, Wenyuan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918163723059200
author Li, Xinfeng
Pang, Shengyuan
Wu, Jialin
Deng, Jiangyi
Zhong, Huanlong
Chen, Yanjiao
Zhang, Jie
Xu, Wenyuan
author_facet Li, Xinfeng
Pang, Shengyuan
Wu, Jialin
Deng, Jiangyi
Zhong, Huanlong
Chen, Yanjiao
Zhang, Jie
Xu, Wenyuan
contents Text-to-image (T2I) models, though exhibiting remarkable creativity in image generation, can be exploited to produce unsafe images. Existing safety measures, e.g., content moderation or model alignment, fail in the presence of white-box adversaries who know and can adjust model parameters, e.g., by fine-tuning. This paper presents a novel defensive framework, named Patronus, which equips T2I models with holistic protection to defend against white-box adversaries. Specifically, we design an internal moderator that decodes unsafe input features into zero vectors while ensuring the decoding performance of benign input features. Furthermore, we strengthen the model alignment with a carefully designed non-fine-tunable learning mechanism, ensuring the T2I model will not be compromised by malicious fine-tuning. We conduct extensive experiments to validate the intactness of the performance on safe content generation and the effectiveness of rejecting unsafe content generation. Results also confirm the resilience of Patronus against various fine-tuning attacks by white-box adversaries.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16581
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Patronus: Safeguarding Text-to-Image Models against White-Box Adversaries
Li, Xinfeng
Pang, Shengyuan
Wu, Jialin
Deng, Jiangyi
Zhong, Huanlong
Chen, Yanjiao
Zhang, Jie
Xu, Wenyuan
Cryptography and Security
Computer Vision and Pattern Recognition
Text-to-image (T2I) models, though exhibiting remarkable creativity in image generation, can be exploited to produce unsafe images. Existing safety measures, e.g., content moderation or model alignment, fail in the presence of white-box adversaries who know and can adjust model parameters, e.g., by fine-tuning. This paper presents a novel defensive framework, named Patronus, which equips T2I models with holistic protection to defend against white-box adversaries. Specifically, we design an internal moderator that decodes unsafe input features into zero vectors while ensuring the decoding performance of benign input features. Furthermore, we strengthen the model alignment with a carefully designed non-fine-tunable learning mechanism, ensuring the T2I model will not be compromised by malicious fine-tuning. We conduct extensive experiments to validate the intactness of the performance on safe content generation and the effectiveness of rejecting unsafe content generation. Results also confirm the resilience of Patronus against various fine-tuning attacks by white-box adversaries.
title Patronus: Safeguarding Text-to-Image Models against White-Box Adversaries
topic Cryptography and Security
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.16581