Multimodal Markup Document Models for Graphic Design Completion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kikuchi, Kotaro, Honda, Ukyo, Inoue, Naoto, Otani, Mayu, Simo-Serra, Edgar, Yamaguchi, Kota
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908691871039488
author Kikuchi, Kotaro
Honda, Ukyo
Inoue, Naoto
Otani, Mayu
Simo-Serra, Edgar
Yamaguchi, Kota
author_facet Kikuchi, Kotaro
Honda, Ukyo
Inoue, Naoto
Otani, Mayu
Simo-Serra, Edgar
Yamaguchi, Kota
contents We introduce MarkupDM, a multimodal markup document model that represents graphic design as an interleaved multimodal document consisting of both markup language and images. Unlike existing holistic approaches that rely on an element-by-attribute grid representation, our representation accommodates variable-length elements, type-dependent attributes, and text content. Inspired by fill-in-the-middle training in code generation, we train the model to complete the missing part of a design document from its surrounding context, allowing it to treat various design tasks in a unified manner. Our model also supports image generation by predicting discrete image tokens through a specialized tokenizer with support for image transparency. We evaluate MarkupDM on three tasks, attribute value, image, and text completion, and demonstrate that it can produce plausible designs consistent with the given context. To further illustrate the flexibility of our approach, we evaluate our approach on a new instruction-guided design completion task where our instruction-tuned MarkupDM compares favorably to state-of-the-art image editing models, especially in textual completion. These findings suggest that multimodal language models with our document representation can serve as a versatile foundation for broad design automation.
format Preprint
id arxiv_https___arxiv_org_abs_2409_19051
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multimodal Markup Document Models for Graphic Design Completion
Kikuchi, Kotaro
Honda, Ukyo
Inoue, Naoto
Otani, Mayu
Simo-Serra, Edgar
Yamaguchi, Kota
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
We introduce MarkupDM, a multimodal markup document model that represents graphic design as an interleaved multimodal document consisting of both markup language and images. Unlike existing holistic approaches that rely on an element-by-attribute grid representation, our representation accommodates variable-length elements, type-dependent attributes, and text content. Inspired by fill-in-the-middle training in code generation, we train the model to complete the missing part of a design document from its surrounding context, allowing it to treat various design tasks in a unified manner. Our model also supports image generation by predicting discrete image tokens through a specialized tokenizer with support for image transparency. We evaluate MarkupDM on three tasks, attribute value, image, and text completion, and demonstrate that it can produce plausible designs consistent with the given context. To further illustrate the flexibility of our approach, we evaluate our approach on a new instruction-guided design completion task where our instruction-tuned MarkupDM compares favorably to state-of-the-art image editing models, especially in textual completion. These findings suggest that multimodal language models with our document representation can serve as a versatile foundation for broad design automation.
title Multimodal Markup Document Models for Graphic Design Completion
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2409.19051