CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic Design

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Hui, Hong, Dexiang, Yang, Maoke, Cheng, Yutao, Zhang, Zhao, Chen, Weidong, Shao, Jie, Wu, Xinglong, Wu, Zuxuan, Jiang, Yu-Gang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917373785669632
author Zhang, Hui
Hong, Dexiang
Yang, Maoke
Cheng, Yutao
Zhang, Zhao
Chen, Weidong
Shao, Jie
Wu, Xinglong
Wu, Zuxuan
Jiang, Yu-Gang
author_facet Zhang, Hui
Hong, Dexiang
Yang, Maoke
Cheng, Yutao
Zhang, Zhao
Chen, Weidong
Shao, Jie
Wu, Xinglong
Wu, Zuxuan
Jiang, Yu-Gang
contents Graphic design plays a vital role in visual communication across advertising, marketing, and multimedia entertainment. Prior work has explored automated graphic design generation using diffusion models, aiming to streamline creative workflows and democratize design capabilities. However, complex graphic design scenarios require accurately adhering to design intent specified by multiple heterogeneous user-provided elements (\eg images, layouts, and texts), which pose multi-condition control challenges for existing methods. Specifically, previous single-condition control models demonstrate effectiveness only within their specialized domains but fail to generalize to other conditions, while existing multi-condition methods often lack fine-grained control over each sub-condition and compromise overall compositional harmony. To address these limitations, we introduce CreatiDesign, a systematic solution for automated graphic design covering both model architecture and dataset construction. First, we design a unified multi-condition driven architecture that enables flexible and precise integration of heterogeneous design elements with minimal architectural modifications to the base diffusion model. Furthermore, to ensure that each condition precisely controls its designated image region and to avoid interference between conditions, we propose a multimodal attention mask mechanism. Additionally, we develop a fully automated pipeline for constructing graphic design datasets, and introduce a new dataset with 400K samples featuring multi-condition annotations, along with a comprehensive benchmark. Experimental results show that CreatiDesign outperforms existing models by a clear margin in faithfully adhering to user intent.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19114
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic Design
Zhang, Hui
Hong, Dexiang
Yang, Maoke
Cheng, Yutao
Zhang, Zhao
Chen, Weidong
Shao, Jie
Wu, Xinglong
Wu, Zuxuan
Jiang, Yu-Gang
Computer Vision and Pattern Recognition
Graphic design plays a vital role in visual communication across advertising, marketing, and multimedia entertainment. Prior work has explored automated graphic design generation using diffusion models, aiming to streamline creative workflows and democratize design capabilities. However, complex graphic design scenarios require accurately adhering to design intent specified by multiple heterogeneous user-provided elements (\eg images, layouts, and texts), which pose multi-condition control challenges for existing methods. Specifically, previous single-condition control models demonstrate effectiveness only within their specialized domains but fail to generalize to other conditions, while existing multi-condition methods often lack fine-grained control over each sub-condition and compromise overall compositional harmony. To address these limitations, we introduce CreatiDesign, a systematic solution for automated graphic design covering both model architecture and dataset construction. First, we design a unified multi-condition driven architecture that enables flexible and precise integration of heterogeneous design elements with minimal architectural modifications to the base diffusion model. Furthermore, to ensure that each condition precisely controls its designated image region and to avoid interference between conditions, we propose a multimodal attention mask mechanism. Additionally, we develop a fully automated pipeline for constructing graphic design datasets, and introduce a new dataset with 400K samples featuring multi-condition annotations, along with a comprehensive benchmark. Experimental results show that CreatiDesign outperforms existing models by a clear margin in faithfully adhering to user intent.
title CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic Design
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.19114