Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Na, Byeonghu, Kang, Mina, Kwak, Jiseok, Park, Minsang, Shin, Jiwoo, Jun, SeJoon, Lee, Gayoung, Kim, Jin-Hwa, Moon, Il-Chul
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909873538596864
author Na, Byeonghu
Kang, Mina
Kwak, Jiseok
Park, Minsang
Shin, Jiwoo
Jun, SeJoon
Lee, Gayoung
Kim, Jin-Hwa
Moon, Il-Chul
author_facet Na, Byeonghu
Kang, Mina
Kwak, Jiseok
Park, Minsang
Shin, Jiwoo
Jun, SeJoon
Lee, Gayoung
Kim, Jin-Hwa
Moon, Il-Chul
contents Text-to-image models have recently made significant advances in generating realistic and semantically coherent images, driven by advanced diffusion models and large-scale web-crawled datasets. However, these datasets often contain inappropriate or biased content, raising concerns about the generation of harmful outputs when provided with malicious text prompts. We propose Safe Text embedding Guidance (STG), a training-free approach to improve the safety of diffusion models by guiding the text embeddings during sampling. STG adjusts the text embeddings based on a safety function evaluated on the expected final denoised image, allowing the model to generate safer outputs without additional training. Theoretically, we show that STG aligns the underlying model distribution with safety constraints, thereby achieving safer outputs while minimally affecting generation quality. Experiments on various safety scenarios, including nudity, violence, and artist-style removal, show that STG consistently outperforms both training-based and training-free baselines in removing unsafe content while preserving the core semantic intent of input prompts. Our code is available at https://github.com/aailab-kaist/STG.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24012
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion Models
Na, Byeonghu
Kang, Mina
Kwak, Jiseok
Park, Minsang
Shin, Jiwoo
Jun, SeJoon
Lee, Gayoung
Kim, Jin-Hwa
Moon, Il-Chul
Machine Learning
Artificial Intelligence
Text-to-image models have recently made significant advances in generating realistic and semantically coherent images, driven by advanced diffusion models and large-scale web-crawled datasets. However, these datasets often contain inappropriate or biased content, raising concerns about the generation of harmful outputs when provided with malicious text prompts. We propose Safe Text embedding Guidance (STG), a training-free approach to improve the safety of diffusion models by guiding the text embeddings during sampling. STG adjusts the text embeddings based on a safety function evaluated on the expected final denoised image, allowing the model to generate safer outputs without additional training. Theoretically, we show that STG aligns the underlying model distribution with safety constraints, thereby achieving safer outputs while minimally affecting generation quality. Experiments on various safety scenarios, including nudity, violence, and artist-style removal, show that STG consistently outperforms both training-based and training-free baselines in removing unsafe content while preserving the core semantic intent of input prompts. Our code is available at https://github.com/aailab-kaist/STG.
title Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.24012