Sketch-to-Layout: Sketch-Guided Multimodal Layout Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Brioschi, Riccardo, Alekseev, Aleksandr, Nevali, Emanuele, Döner, Berkay, Malki, Omar El, Mitrevski, Blagoj, Kieliger, Leandro, Collier, Mark, Maksai, Andrii, Berent, Jesse, Musat, Claudiu, Kokiopoulou, Efi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911243675107328
author Brioschi, Riccardo
Alekseev, Aleksandr
Nevali, Emanuele
Döner, Berkay
Malki, Omar El
Mitrevski, Blagoj
Kieliger, Leandro
Collier, Mark
Maksai, Andrii
Berent, Jesse
Musat, Claudiu
Kokiopoulou, Efi
author_facet Brioschi, Riccardo
Alekseev, Aleksandr
Nevali, Emanuele
Döner, Berkay
Malki, Omar El
Mitrevski, Blagoj
Kieliger, Leandro
Collier, Mark
Maksai, Andrii
Berent, Jesse
Musat, Claudiu
Kokiopoulou, Efi
contents Graphic layout generation is a growing research area focusing on generating aesthetically pleasing layouts ranging from poster designs to documents. While recent research has explored ways to incorporate user constraints to guide the layout generation, these constraints often require complex specifications which reduce usability. We introduce an innovative approach exploiting user-provided sketches as intuitive constraints and we demonstrate empirically the effectiveness of this new guidance method, establishing the sketch-to-layout problem as a promising research direction, which is currently under-explored. To tackle the sketch-to-layout problem, we propose a multimodal transformer-based solution using the sketch and the content assets as inputs to produce high quality layouts. Since collecting sketch training data from human annotators to train our model is very costly, we introduce a novel and efficient method to synthetically generate training sketches at scale. We train and evaluate our model on three publicly available datasets: PubLayNet, DocLayNet and SlidesVQA, demonstrating that it outperforms state-of-the-art constraint-based methods, while offering a more intuitive design experience. In order to facilitate future sketch-to-layout research, we release O(200k) synthetically-generated sketches for the public datasets above. The datasets are available at https://github.com/google-deepmind/sketch_to_layout.
format Preprint
id arxiv_https___arxiv_org_abs_2510_27632
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sketch-to-Layout: Sketch-Guided Multimodal Layout Generation
Brioschi, Riccardo
Alekseev, Aleksandr
Nevali, Emanuele
Döner, Berkay
Malki, Omar El
Mitrevski, Blagoj
Kieliger, Leandro
Collier, Mark
Maksai, Andrii
Berent, Jesse
Musat, Claudiu
Kokiopoulou, Efi
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphic layout generation is a growing research area focusing on generating aesthetically pleasing layouts ranging from poster designs to documents. While recent research has explored ways to incorporate user constraints to guide the layout generation, these constraints often require complex specifications which reduce usability. We introduce an innovative approach exploiting user-provided sketches as intuitive constraints and we demonstrate empirically the effectiveness of this new guidance method, establishing the sketch-to-layout problem as a promising research direction, which is currently under-explored. To tackle the sketch-to-layout problem, we propose a multimodal transformer-based solution using the sketch and the content assets as inputs to produce high quality layouts. Since collecting sketch training data from human annotators to train our model is very costly, we introduce a novel and efficient method to synthetically generate training sketches at scale. We train and evaluate our model on three publicly available datasets: PubLayNet, DocLayNet and SlidesVQA, demonstrating that it outperforms state-of-the-art constraint-based methods, while offering a more intuitive design experience. In order to facilitate future sketch-to-layout research, we release O(200k) synthetically-generated sketches for the public datasets above. The datasets are available at https://github.com/google-deepmind/sketch_to_layout.
title Sketch-to-Layout: Sketch-Guided Multimodal Layout Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.27632