Be Decisive: Noise-Induced Layouts for Multi-Subject Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dahary, Omer, Cohen, Yehonathan, Patashnik, Or, Aberman, Kfir, Cohen-Or, Daniel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916762053771264
author Dahary, Omer
Cohen, Yehonathan
Patashnik, Or
Aberman, Kfir
Cohen-Or, Daniel
author_facet Dahary, Omer
Cohen, Yehonathan
Patashnik, Or
Aberman, Kfir
Cohen-Or, Daniel
contents Generating multiple distinct subjects remains a challenge for existing text-to-image diffusion models. Complex prompts often lead to subject leakage, causing inaccuracies in quantities, attributes, and visual features. Preventing leakage among subjects necessitates knowledge of each subject's spatial location. Recent methods provide these spatial locations via an external layout control. However, enforcing such a prescribed layout often conflicts with the innate layout dictated by the sampled initial noise, leading to misalignment with the model's prior. In this work, we introduce a new approach that predicts a spatial layout aligned with the prompt, derived from the initial noise, and refines it throughout the denoising process. By relying on this noise-induced layout, we avoid conflicts with externally imposed layouts and better preserve the model's prior. Our method employs a small neural network to predict and refine the evolving noise-induced layout at each denoising step, ensuring clear boundaries between subjects while maintaining consistency. Experimental results show that this noise-aligned strategy achieves improved text-image alignment and more stable multi-subject generation compared to existing layout-guided techniques, while preserving the rich diversity of the model's original distribution.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21488
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Be Decisive: Noise-Induced Layouts for Multi-Subject Generation
Dahary, Omer
Cohen, Yehonathan
Patashnik, Or
Aberman, Kfir
Cohen-Or, Daniel
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
Generating multiple distinct subjects remains a challenge for existing text-to-image diffusion models. Complex prompts often lead to subject leakage, causing inaccuracies in quantities, attributes, and visual features. Preventing leakage among subjects necessitates knowledge of each subject's spatial location. Recent methods provide these spatial locations via an external layout control. However, enforcing such a prescribed layout often conflicts with the innate layout dictated by the sampled initial noise, leading to misalignment with the model's prior. In this work, we introduce a new approach that predicts a spatial layout aligned with the prompt, derived from the initial noise, and refines it throughout the denoising process. By relying on this noise-induced layout, we avoid conflicts with externally imposed layouts and better preserve the model's prior. Our method employs a small neural network to predict and refine the evolving noise-induced layout at each denoising step, ensuring clear boundaries between subjects while maintaining consistency. Experimental results show that this noise-aligned strategy achieves improved text-image alignment and more stable multi-subject generation compared to existing layout-guided techniques, while preserving the rich diversity of the model's original distribution.
title Be Decisive: Noise-Induced Layouts for Multi-Subject Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
url https://arxiv.org/abs/2505.21488