SLayR: Scene Layout Generation with Rectified Flow

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Braunstein, Cameron, Petekkaya, Hevra, Lenssen, Jan Eric, Toneva, Mariya, Ilg, Eddy
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915193535070208
author Braunstein, Cameron
Petekkaya, Hevra
Lenssen, Jan Eric
Toneva, Mariya
Ilg, Eddy
author_facet Braunstein, Cameron
Petekkaya, Hevra
Lenssen, Jan Eric
Toneva, Mariya
Ilg, Eddy
contents We introduce SLayR, Scene Layout Generation with Rectified flow, a novel transformer-based model for text-to-layout generation which can then be paired with existing layout-to-image models to produce images. SLayR addresses a domain in which current text-to-image pipelines struggle: generating scene layouts that are of significant variety and plausibility, when the given prompt is ambiguous and does not provide constraints on the scene. SLayR surpasses existing baselines including LLMs in unconstrained generation, and can generate layouts from an open caption set. To accurately evaluate the layout generation, we introduce a new benchmark suite, including numerical metrics and a carefully designed repeatable human-evaluation procedure that assesses the plausibility and variety of generated images. We show that our method sets a new state of the art for achieving both at the same time, while being at least 3x times smaller in the number of parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2412_05003
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SLayR: Scene Layout Generation with Rectified Flow
Braunstein, Cameron
Petekkaya, Hevra
Lenssen, Jan Eric
Toneva, Mariya
Ilg, Eddy
Computer Vision and Pattern Recognition
We introduce SLayR, Scene Layout Generation with Rectified flow, a novel transformer-based model for text-to-layout generation which can then be paired with existing layout-to-image models to produce images. SLayR addresses a domain in which current text-to-image pipelines struggle: generating scene layouts that are of significant variety and plausibility, when the given prompt is ambiguous and does not provide constraints on the scene. SLayR surpasses existing baselines including LLMs in unconstrained generation, and can generate layouts from an open caption set. To accurately evaluate the layout generation, we introduce a new benchmark suite, including numerical metrics and a carefully designed repeatable human-evaluation procedure that assesses the plausibility and variety of generated images. We show that our method sets a new state of the art for achieving both at the same time, while being at least 3x times smaller in the number of parameters.
title SLayR: Scene Layout Generation with Rectified Flow
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.05003