Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Yutong, Gong, Biao, Chen, Di, Shen, Yujun, Liu, Yu, Zhou, Jingren
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916198147424256
author Feng, Yutong
Gong, Biao
Chen, Di
Shen, Yujun
Liu, Yu
Zhou, Jingren
author_facet Feng, Yutong
Gong, Biao
Chen, Di
Shen, Yujun
Liu, Yu
Zhou, Jingren
contents Existing text-to-image (T2I) diffusion models usually struggle in interpreting complex prompts, especially those with quantity, object-attribute binding, and multi-subject descriptions. In this work, we introduce a semantic panel as the middleware in decoding texts to images, supporting the generator to better follow instructions. The panel is obtained through arranging the visual concepts parsed from the input text by the aid of large language models, and then injected into the denoising network as a detailed control signal to complement the text condition. To facilitate text-to-panel learning, we come up with a carefully designed semantic formatting protocol, accompanied by a fully-automatic data preparation pipeline. Thanks to such a design, our approach, which we call Ranni, manages to enhance a pre-trained T2I generator regarding its textual controllability. More importantly, the introduction of the generative middleware brings a more convenient form of interaction (i.e., directly adjusting the elements in the panel or using language instructions) and further allows users to finely customize their generation, based on which we develop a practical system and showcase its potential in continuous generation and chatting-based editing. Our project page is at https://ranni-t2i.github.io/Ranni.
format Preprint
id arxiv_https___arxiv_org_abs_2311_17002
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following
Feng, Yutong
Gong, Biao
Chen, Di
Shen, Yujun
Liu, Yu
Zhou, Jingren
Computer Vision and Pattern Recognition
Existing text-to-image (T2I) diffusion models usually struggle in interpreting complex prompts, especially those with quantity, object-attribute binding, and multi-subject descriptions. In this work, we introduce a semantic panel as the middleware in decoding texts to images, supporting the generator to better follow instructions. The panel is obtained through arranging the visual concepts parsed from the input text by the aid of large language models, and then injected into the denoising network as a detailed control signal to complement the text condition. To facilitate text-to-panel learning, we come up with a carefully designed semantic formatting protocol, accompanied by a fully-automatic data preparation pipeline. Thanks to such a design, our approach, which we call Ranni, manages to enhance a pre-trained T2I generator regarding its textual controllability. More importantly, the introduction of the generative middleware brings a more convenient form of interaction (i.e., directly adjusting the elements in the panel or using language instructions) and further allows users to finely customize their generation, based on which we develop a practical system and showcase its potential in continuous generation and chatting-based editing. Our project page is at https://ranni-t2i.github.io/Ranni.
title Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.17002