VersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Zhipeng, Yang, Lan, Qi, Yonggang, Zhang, Honggang, Pang, Kaiyue, Li, Ke, Song, Yi-Zhe
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915083267866624
author Chen, Zhipeng
Yang, Lan
Qi, Yonggang
Zhang, Honggang
Pang, Kaiyue
Li, Ke
Song, Yi-Zhe
author_facet Chen, Zhipeng
Yang, Lan
Qi, Yonggang
Zhang, Honggang
Pang, Kaiyue
Li, Ke
Song, Yi-Zhe
contents Despite the rapid advancements in text-to-image (T2I) synthesis, enabling precise visual control remains a significant challenge. Existing works attempted to incorporate multi-facet controls (text and sketch), aiming to enhance the creative control over generated images. However, our pilot study reveals that the expressive power of humans far surpasses the capabilities of current methods. Users desire a more versatile approach that can accommodate their diverse creative intents, ranging from controlling individual subjects to manipulating the entire scene composition. We present VersaGen, a generative AI agent that enables versatile visual control in T2I synthesis. VersaGen admits four types of visual controls: i) single visual subject; ii) multiple visual subjects; iii) scene background; iv) any combination of the three above or merely no control at all. We train an adaptor upon a frozen T2I model to accommodate the visual information into the text-dominated diffusion process. We introduce three optimization strategies during the inference phase of VersaGen to improve generation results and enhance user experience. Comprehensive experiments on COCO and Sketchy validate the effectiveness and flexibility of VersaGen, as evidenced by both qualitative and quantitative results.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11594
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis
Chen, Zhipeng
Yang, Lan
Qi, Yonggang
Zhang, Honggang
Pang, Kaiyue
Li, Ke
Song, Yi-Zhe
Computer Vision and Pattern Recognition
I.4.9; I.4.10
Despite the rapid advancements in text-to-image (T2I) synthesis, enabling precise visual control remains a significant challenge. Existing works attempted to incorporate multi-facet controls (text and sketch), aiming to enhance the creative control over generated images. However, our pilot study reveals that the expressive power of humans far surpasses the capabilities of current methods. Users desire a more versatile approach that can accommodate their diverse creative intents, ranging from controlling individual subjects to manipulating the entire scene composition. We present VersaGen, a generative AI agent that enables versatile visual control in T2I synthesis. VersaGen admits four types of visual controls: i) single visual subject; ii) multiple visual subjects; iii) scene background; iv) any combination of the three above or merely no control at all. We train an adaptor upon a frozen T2I model to accommodate the visual information into the text-dominated diffusion process. We introduce three optimization strategies during the inference phase of VersaGen to improve generation results and enhance user experience. Comprehensive experiments on COCO and Sketchy validate the effectiveness and flexibility of VersaGen, as evidenced by both qualitative and quantitative results.
title VersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis
topic Computer Vision and Pattern Recognition
I.4.9; I.4.10
url https://arxiv.org/abs/2412.11594