Saved in:
Bibliographic Details
Main Authors: Chen, Jingye, Wang, Zhaowen, Zhao, Nanxuan, Zhang, Li, Liu, Difan, Yang, Jimei, Chen, Qifeng
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.05601
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915376769531904
author Chen, Jingye
Wang, Zhaowen
Zhao, Nanxuan
Zhang, Li
Liu, Difan
Yang, Jimei
Chen, Qifeng
author_facet Chen, Jingye
Wang, Zhaowen
Zhao, Nanxuan
Zhang, Li
Liu, Difan
Yang, Jimei
Chen, Qifeng
contents Graphic design is crucial for conveying ideas and messages. Designers usually organize their work into objects, backgrounds, and vectorized text layers to simplify editing. However, this workflow demands considerable expertise. With the rise of GenAI methods, an endless supply of high-quality graphic designs in pixel format has become more accessible, though these designs often lack editability. Despite this, non-layered designs still inspire human designers, influencing their choices in layouts and text styles, ultimately guiding the creation of layered designs. Motivated by this observation, we propose Accordion, a graphic design generation framework taking the first attempt to convert AI-generated designs into editable layered designs, meanwhile refining nonsensical AI-generated text with meaningful alternatives guided by user prompts. It is built around a vision language model (VLM) playing distinct roles in three curated stages. For each stage, we design prompts to guide the VLM in executing different tasks. Distinct from existing bottom-up methods (e.g., COLE and Open-COLE) that gradually generate elements to create layered designs, our approach works in a top-down manner by using the visually harmonious reference image as global guidance to decompose each layer. Additionally, it leverages multiple vision experts such as SAM and element removal models to facilitate the creation of graphic layers. We train our method using the in-house graphic design dataset Design39K, augmented with AI-generated design images coupled with refined ground truth created by a customized inpainting model. Experimental results and user studies by designers show that Accordion generates favorable results on the DesignIntention benchmark, including tasks such as text-to-template, adding text to background, and text de-rendering, and also excels in creating design variations.
format Preprint
id arxiv_https___arxiv_org_abs_2507_05601
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking Layered Graphic Design Generation with a Top-Down Approach
Chen, Jingye
Wang, Zhaowen
Zhao, Nanxuan
Zhang, Li
Liu, Difan
Yang, Jimei
Chen, Qifeng
Computer Vision and Pattern Recognition
Graphic design is crucial for conveying ideas and messages. Designers usually organize their work into objects, backgrounds, and vectorized text layers to simplify editing. However, this workflow demands considerable expertise. With the rise of GenAI methods, an endless supply of high-quality graphic designs in pixel format has become more accessible, though these designs often lack editability. Despite this, non-layered designs still inspire human designers, influencing their choices in layouts and text styles, ultimately guiding the creation of layered designs. Motivated by this observation, we propose Accordion, a graphic design generation framework taking the first attempt to convert AI-generated designs into editable layered designs, meanwhile refining nonsensical AI-generated text with meaningful alternatives guided by user prompts. It is built around a vision language model (VLM) playing distinct roles in three curated stages. For each stage, we design prompts to guide the VLM in executing different tasks. Distinct from existing bottom-up methods (e.g., COLE and Open-COLE) that gradually generate elements to create layered designs, our approach works in a top-down manner by using the visually harmonious reference image as global guidance to decompose each layer. Additionally, it leverages multiple vision experts such as SAM and element removal models to facilitate the creation of graphic layers. We train our method using the in-house graphic design dataset Design39K, augmented with AI-generated design images coupled with refined ground truth created by a customized inpainting model. Experimental results and user studies by designers show that Accordion generates favorable results on the DesignIntention benchmark, including tasks such as text-to-template, adding text to background, and text de-rendering, and also excels in creating design variations.
title Rethinking Layered Graphic Design Generation with a Top-Down Approach
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.05601