Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Baumann, Stefan Andreas, Krause, Felix, Neumayr, Michael, Stracke, Nick, Sevi, Melvin, Hu, Vincent Tao, Ommer, Björn
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917956353523712
author Baumann, Stefan Andreas
Krause, Felix
Neumayr, Michael
Stracke, Nick
Sevi, Melvin
Hu, Vincent Tao
Ommer, Björn
author_facet Baumann, Stefan Andreas
Krause, Felix
Neumayr, Michael
Stracke, Nick
Sevi, Melvin
Hu, Vincent Tao
Ommer, Björn
contents Recent advances in text-to-image (T2I) diffusion models have significantly improved the quality of generated images. However, providing efficient control over individual subjects, particularly the attributes characterizing them, remains a key challenge. While existing methods have introduced mechanisms to modulate attribute expression, they typically provide either detailed, object-specific localization of such a modification or full-scale fine-grained, nuanced control of attributes. No current approach offers both simultaneously, resulting in a gap when trying to achieve precise continuous and subject-specific attribute modulation in image generation. In this work, we demonstrate that token-level directions exist within commonly used CLIP text embeddings that enable fine-grained, subject-specific control of high-level attributes in T2I models. We introduce two methods to identify these directions: a simple, optimization-free technique and a learning-based approach that utilizes the T2I model to characterize semantic concepts more specifically. Our methods allow the augmentation of the prompt text input, enabling fine-grained control over multiple attributes of individual subjects simultaneously, without requiring any modifications to the diffusion model itself. This approach offers a unified solution that fills the gap between global and localized control, providing competitive flexibility and precision in text-guided image generation. Project page: https://compvis.github.io/attribute-control. Code is available at https://github.com/CompVis/attribute-control.
format Preprint
id arxiv_https___arxiv_org_abs_2403_17064
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions
Baumann, Stefan Andreas
Krause, Felix
Neumayr, Michael
Stracke, Nick
Sevi, Melvin
Hu, Vincent Tao
Ommer, Björn
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Recent advances in text-to-image (T2I) diffusion models have significantly improved the quality of generated images. However, providing efficient control over individual subjects, particularly the attributes characterizing them, remains a key challenge. While existing methods have introduced mechanisms to modulate attribute expression, they typically provide either detailed, object-specific localization of such a modification or full-scale fine-grained, nuanced control of attributes. No current approach offers both simultaneously, resulting in a gap when trying to achieve precise continuous and subject-specific attribute modulation in image generation. In this work, we demonstrate that token-level directions exist within commonly used CLIP text embeddings that enable fine-grained, subject-specific control of high-level attributes in T2I models. We introduce two methods to identify these directions: a simple, optimization-free technique and a learning-based approach that utilizes the T2I model to characterize semantic concepts more specifically. Our methods allow the augmentation of the prompt text input, enabling fine-grained control over multiple attributes of individual subjects simultaneously, without requiring any modifications to the diffusion model itself. This approach offers a unified solution that fills the gap between global and localized control, providing competitive flexibility and precision in text-guided image generation. Project page: https://compvis.github.io/attribute-control. Code is available at https://github.com/CompVis/attribute-control.
title Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2403.17064