Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Ruyu, Hou, Xuefeng, Schmedding, Sabrina, Huber, Marco F.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2503.12213
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917958548193280
author Wang, Ruyu
Hou, Xuefeng
Schmedding, Sabrina
Huber, Marco F.
author_facet Wang, Ruyu
Hou, Xuefeng
Schmedding, Sabrina
Huber, Marco F.
contents In layout-to-image (L2I) synthesis, controlled complex scenes are generated from coarse information like bounding boxes. Such a task is exciting to many downstream applications because the input layouts offer strong guidance to the generation process while remaining easily reconfigurable by humans. In this paper, we proposed STyled LAYout Diffusion (STAY Diffusion), a diffusion-based model that produces photo-realistic images and provides fine-grained control of stylized objects in scenes. Our approach learns a global condition for each layout, and a self-supervised semantic map for weight modulation using a novel Edge-Aware Normalization (EA Norm). A new Styled-Mask Attention (SM Attention) is also introduced to cross-condition the global condition and image feature for capturing the objects' relationships. These measures provide consistent guidance through the model, enabling more accurate and controllable image generation. Extensive benchmarking demonstrates that our STAY Diffusion presents high-quality images while surpassing previous state-of-the-art methods in generation diversity, accuracy, and controllability.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12213
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle STAY Diffusion: Styled Layout Diffusion Model for Diverse Layout-to-Image Generation
Wang, Ruyu
Hou, Xuefeng
Schmedding, Sabrina
Huber, Marco F.
Computer Vision and Pattern Recognition
In layout-to-image (L2I) synthesis, controlled complex scenes are generated from coarse information like bounding boxes. Such a task is exciting to many downstream applications because the input layouts offer strong guidance to the generation process while remaining easily reconfigurable by humans. In this paper, we proposed STyled LAYout Diffusion (STAY Diffusion), a diffusion-based model that produces photo-realistic images and provides fine-grained control of stylized objects in scenes. Our approach learns a global condition for each layout, and a self-supervised semantic map for weight modulation using a novel Edge-Aware Normalization (EA Norm). A new Styled-Mask Attention (SM Attention) is also introduced to cross-condition the global condition and image feature for capturing the objects' relationships. These measures provide consistent guidance through the model, enabling more accurate and controllable image generation. Extensive benchmarking demonstrates that our STAY Diffusion presents high-quality images while surpassing previous state-of-the-art methods in generation diversity, accuracy, and controllability.
title STAY Diffusion: Styled Layout Diffusion Model for Diverse Layout-to-Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.12213