Self-Discovering Interpretable Diffusion Latent Directions for Responsible Text-to-Image Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Hang, Shen, Chengzhi, Torr, Philip, Tresp, Volker, Gu, Jindong
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914731389878272
author Li, Hang
Shen, Chengzhi
Torr, Philip
Tresp, Volker
Gu, Jindong
author_facet Li, Hang
Shen, Chengzhi
Torr, Philip
Tresp, Volker
Gu, Jindong
contents Diffusion-based models have gained significant popularity for text-to-image generation due to their exceptional image-generation capabilities. A risk with these models is the potential generation of inappropriate content, such as biased or harmful images. However, the underlying reasons for generating such undesired content from the perspective of the diffusion model's internal representation remain unclear. Previous work interprets vectors in an interpretable latent space of diffusion models as semantic concepts. However, existing approaches cannot discover directions for arbitrary concepts, such as those related to inappropriate concepts. In this work, we propose a novel self-supervised approach to find interpretable latent directions for a given concept. With the discovered vectors, we further propose a simple approach to mitigate inappropriate generation. Extensive experiments have been conducted to verify the effectiveness of our mitigation approach, namely, for fair generation, safe generation, and responsible text-enhancing generation. Project page: \url{https://interpretdiffusion.github.io}.
format Preprint
id arxiv_https___arxiv_org_abs_2311_17216
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Self-Discovering Interpretable Diffusion Latent Directions for Responsible Text-to-Image Generation
Li, Hang
Shen, Chengzhi
Torr, Philip
Tresp, Volker
Gu, Jindong
Computer Vision and Pattern Recognition
Diffusion-based models have gained significant popularity for text-to-image generation due to their exceptional image-generation capabilities. A risk with these models is the potential generation of inappropriate content, such as biased or harmful images. However, the underlying reasons for generating such undesired content from the perspective of the diffusion model's internal representation remain unclear. Previous work interprets vectors in an interpretable latent space of diffusion models as semantic concepts. However, existing approaches cannot discover directions for arbitrary concepts, such as those related to inappropriate concepts. In this work, we propose a novel self-supervised approach to find interpretable latent directions for a given concept. With the discovered vectors, we further propose a simple approach to mitigate inappropriate generation. Extensive experiments have been conducted to verify the effectiveness of our mitigation approach, namely, for fair generation, safe generation, and responsible text-enhancing generation. Project page: \url{https://interpretdiffusion.github.io}.
title Self-Discovering Interpretable Diffusion Latent Directions for Responsible Text-to-Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.17216