TokenCompose: Text-to-Image Diffusion with Token-level Supervision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zirui, Sha, Zhizhou, Ding, Zheng, Wang, Yilin, Tu, Zhuowen
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913402223329280
author Wang, Zirui
Sha, Zhizhou
Ding, Zheng
Wang, Yilin
Tu, Zhuowen
author_facet Wang, Zirui
Sha, Zhizhou
Ding, Zheng
Wang, Yilin
Tu, Zhuowen
contents We present TokenCompose, a Latent Diffusion Model for text-to-image generation that achieves enhanced consistency between user-specified text prompts and model-generated images. Despite its tremendous success, the standard denoising process in the Latent Diffusion Model takes text prompts as conditions only, absent explicit constraint for the consistency between the text prompts and the image contents, leading to unsatisfactory results for composing multiple object categories. TokenCompose aims to improve multi-category instance composition by introducing the token-wise consistency terms between the image content and object segmentation maps in the finetuning stage. TokenCompose can be applied directly to the existing training pipeline of text-conditioned diffusion models without extra human labeling information. By finetuning Stable Diffusion, the model exhibits significant improvements in multi-category instance composition and enhanced photorealism for its generated images. Project link: https://mlpc-ucsd.github.io/TokenCompose
format Preprint
id arxiv_https___arxiv_org_abs_2312_03626
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle TokenCompose: Text-to-Image Diffusion with Token-level Supervision
Wang, Zirui
Sha, Zhizhou
Ding, Zheng
Wang, Yilin
Tu, Zhuowen
Computer Vision and Pattern Recognition
We present TokenCompose, a Latent Diffusion Model for text-to-image generation that achieves enhanced consistency between user-specified text prompts and model-generated images. Despite its tremendous success, the standard denoising process in the Latent Diffusion Model takes text prompts as conditions only, absent explicit constraint for the consistency between the text prompts and the image contents, leading to unsatisfactory results for composing multiple object categories. TokenCompose aims to improve multi-category instance composition by introducing the token-wise consistency terms between the image content and object segmentation maps in the finetuning stage. TokenCompose can be applied directly to the existing training pipeline of text-conditioned diffusion models without extra human labeling information. By finetuning Stable Diffusion, the model exhibits significant improvements in multi-category instance composition and enhanced photorealism for its generated images. Project link: https://mlpc-ucsd.github.io/TokenCompose
title TokenCompose: Text-to-Image Diffusion with Token-level Supervision
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.03626