DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Dongzhi, Zhang, Renrui, Li, Haodong, Zong, Zhuofan, Guo, Ziyu, He, Jun, Guo, Claire, Ye, Junyan, Fang, Rongyao, Li, Weijia, Liu, Rui, Li, Hongsheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918232102797312
author Jiang, Dongzhi
Zhang, Renrui
Li, Haodong
Zong, Zhuofan
Guo, Ziyu
He, Jun
Guo, Claire
Ye, Junyan
Fang, Rongyao
Li, Weijia
Liu, Rui
Li, Hongsheng
author_facet Jiang, Dongzhi
Zhang, Renrui
Li, Haodong
Zong, Zhuofan
Guo, Ziyu
He, Jun
Guo, Claire
Ye, Junyan
Fang, Rongyao
Li, Weijia
Liu, Rui
Li, Hongsheng
contents Recent unified multimodal large language models (MLLMs) have shown impressive capabilities, incorporating chain-of-thought (CoT) reasoning for enhanced text-to-image generation. However, existing approaches remain limited, either treating the model merely as a standalone generator or relying on abstract textual planning. To this end, we propose Draft-as-CoT (DraCo), a novel interleaved reasoning paradigm that fully leverages both textual and visual contents in CoT for better planning and verification. Our method first generates a low-resolution draft image as preview, providing more concrete and structural visual planning and guidance. Then, we employ the model's inherent understanding capability to verify potential semantic misalignments between the draft and input prompt, and performs refinement through selective corrections with super-resolution. In this way, our approach addresses two fundamental challenges: the coarse-grained nature of textual planning and the difficulty in generating rare attribute combinations. To support training, we curate DraCo-240K, aiming to enhance three atomic capabilities spanning general correction, instance manipulation, and layout reorganization. Supported by DraCo-CFG, a specialized classifier-free guidance (CFG) strategy for interleaved reasoning, DraCo achieves a tremendous increase on GenEval (+8%), Imagine-Bench (+0.91), and GenEval++ (+3%), significantly outperforming direct generation and other generation methods empowered by CoT.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05112
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
Jiang, Dongzhi
Zhang, Renrui
Li, Haodong
Zong, Zhuofan
Guo, Ziyu
He, Jun
Guo, Claire
Ye, Junyan
Fang, Rongyao
Li, Weijia
Liu, Rui
Li, Hongsheng
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Recent unified multimodal large language models (MLLMs) have shown impressive capabilities, incorporating chain-of-thought (CoT) reasoning for enhanced text-to-image generation. However, existing approaches remain limited, either treating the model merely as a standalone generator or relying on abstract textual planning. To this end, we propose Draft-as-CoT (DraCo), a novel interleaved reasoning paradigm that fully leverages both textual and visual contents in CoT for better planning and verification. Our method first generates a low-resolution draft image as preview, providing more concrete and structural visual planning and guidance. Then, we employ the model's inherent understanding capability to verify potential semantic misalignments between the draft and input prompt, and performs refinement through selective corrections with super-resolution. In this way, our approach addresses two fundamental challenges: the coarse-grained nature of textual planning and the difficulty in generating rare attribute combinations. To support training, we curate DraCo-240K, aiming to enhance three atomic capabilities spanning general correction, instance manipulation, and layout reorganization. Supported by DraCo-CFG, a specialized classifier-free guidance (CFG) strategy for interleaved reasoning, DraCo achieves a tremendous increase on GenEval (+8%), Imagine-Bench (+0.91), and GenEval++ (+3%), significantly outperforming direct generation and other generation methods empowered by CoT.
title DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2512.05112