SceneBooth: Diffusion-based Framework for Subject-preserved Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chai, Shang, Lin, Zihang, Zhou, Min, Li, Xubin, Zhuang, Liansheng, Li, Houqiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916553609445376
author Chai, Shang
Lin, Zihang
Zhou, Min
Li, Xubin
Zhuang, Liansheng
Li, Houqiang
author_facet Chai, Shang
Lin, Zihang
Zhou, Min
Li, Xubin
Zhuang, Liansheng
Li, Houqiang
contents Due to the demand for personalizing image generation, subject-driven text-to-image generation method, which creates novel renditions of an input subject based on text prompts, has received growing research interest. Existing methods often learn subject representation and incorporate it into the prompt embedding to guide image generation, but they struggle with preserving subject fidelity. To solve this issue, this paper approaches a novel framework named SceneBooth for subject-preserved text-to-image generation, which consumes inputs of a subject image, object phrases and text prompts. Instead of learning the subject representation and generating a subject, our SceneBooth fixes the given subject image and generates its background image guided by the text prompts. To this end, our SceneBooth introduces two key components, i.e., a multimodal layout generation module and a background painting module. The former determines the position and scale of the subject by generating appropriate scene layouts that align with text captions, object phrases, and subject visual information. The latter integrates two adapters (ControlNet and Gated Self-Attention) into the latent diffusion model to generate a background that harmonizes with the subject guided by scene layouts and text descriptions. In this manner, our SceneBooth ensures accurate preservation of the subject's appearance in the output. Quantitative and qualitative experimental results demonstrate that SceneBooth significantly outperforms baseline methods in terms of subject preservation, image harmonization and overall quality.
format Preprint
id arxiv_https___arxiv_org_abs_2501_03490
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SceneBooth: Diffusion-based Framework for Subject-preserved Text-to-Image Generation
Chai, Shang
Lin, Zihang
Zhou, Min
Li, Xubin
Zhuang, Liansheng
Li, Houqiang
Computer Vision and Pattern Recognition
Due to the demand for personalizing image generation, subject-driven text-to-image generation method, which creates novel renditions of an input subject based on text prompts, has received growing research interest. Existing methods often learn subject representation and incorporate it into the prompt embedding to guide image generation, but they struggle with preserving subject fidelity. To solve this issue, this paper approaches a novel framework named SceneBooth for subject-preserved text-to-image generation, which consumes inputs of a subject image, object phrases and text prompts. Instead of learning the subject representation and generating a subject, our SceneBooth fixes the given subject image and generates its background image guided by the text prompts. To this end, our SceneBooth introduces two key components, i.e., a multimodal layout generation module and a background painting module. The former determines the position and scale of the subject by generating appropriate scene layouts that align with text captions, object phrases, and subject visual information. The latter integrates two adapters (ControlNet and Gated Self-Attention) into the latent diffusion model to generate a background that harmonizes with the subject guided by scene layouts and text descriptions. In this manner, our SceneBooth ensures accurate preservation of the subject's appearance in the output. Quantitative and qualitative experimental results demonstrate that SceneBooth significantly outperforms baseline methods in terms of subject preservation, image harmonization and overall quality.
title SceneBooth: Diffusion-based Framework for Subject-preserved Text-to-Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.03490