LayerCraft: Enhancing Text-to-Image Generation with CoT Reasoning and Layered Object Integration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yuyao, Li, Jinghao, Tai, Yu-Wing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909851446149120
author Zhang, Yuyao
Li, Jinghao
Tai, Yu-Wing
author_facet Zhang, Yuyao
Li, Jinghao
Tai, Yu-Wing
contents Text-to-image (T2I) generation has made remarkable progress, yet existing systems still lack intuitive control over spatial composition, object consistency, and multi-step editing. We present $\textbf{LayerCraft}$, a modular framework that uses large language models (LLMs) as autonomous agents to orchestrate structured, layered image generation and editing. LayerCraft supports two key capabilities: (1) $\textit{structured generation}$ from simple prompts via chain-of-thought (CoT) reasoning, enabling it to decompose scenes, reason about object placement, and guide composition in a controllable, interpretable manner; and (2) $\textit{layered object integration}$, allowing users to insert and customize objects -- such as characters or props -- across diverse images or scenes while preserving identity, context, and style. The system comprises a coordinator agent, the $\textbf{ChainArchitect}$ for CoT-driven layout planning, and the $\textbf{Object Integration Network (OIN)}$ for seamless image editing using off-the-shelf T2I models without retraining. Through applications like batch collage editing and narrative scene generation, LayerCraft empowers non-experts to iteratively design, customize, and refine visual content with minimal manual effort. Code will be released at https://github.com/PeterYYZhang/LayerCraft.
format Preprint
id arxiv_https___arxiv_org_abs_2504_00010
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LayerCraft: Enhancing Text-to-Image Generation with CoT Reasoning and Layered Object Integration
Zhang, Yuyao
Li, Jinghao
Tai, Yu-Wing
Machine Learning
Graphics
Multiagent Systems
Text-to-image (T2I) generation has made remarkable progress, yet existing systems still lack intuitive control over spatial composition, object consistency, and multi-step editing. We present $\textbf{LayerCraft}$, a modular framework that uses large language models (LLMs) as autonomous agents to orchestrate structured, layered image generation and editing. LayerCraft supports two key capabilities: (1) $\textit{structured generation}$ from simple prompts via chain-of-thought (CoT) reasoning, enabling it to decompose scenes, reason about object placement, and guide composition in a controllable, interpretable manner; and (2) $\textit{layered object integration}$, allowing users to insert and customize objects -- such as characters or props -- across diverse images or scenes while preserving identity, context, and style. The system comprises a coordinator agent, the $\textbf{ChainArchitect}$ for CoT-driven layout planning, and the $\textbf{Object Integration Network (OIN)}$ for seamless image editing using off-the-shelf T2I models without retraining. Through applications like batch collage editing and narrative scene generation, LayerCraft empowers non-experts to iteratively design, customize, and refine visual content with minimal manual effort. Code will be released at https://github.com/PeterYYZhang/LayerCraft.
title LayerCraft: Enhancing Text-to-Image Generation with CoT Reasoning and Layered Object Integration
topic Machine Learning
Graphics
Multiagent Systems
url https://arxiv.org/abs/2504.00010