LACONIC: A 3D Layout Adapter for Controllable Image Creation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Maillard, Léopold, Durand, Tom, Rahary, Adrien Ramanana, Ovsjanikov, Maks
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916877026983936
author Maillard, Léopold
Durand, Tom
Rahary, Adrien Ramanana
Ovsjanikov, Maks
author_facet Maillard, Léopold
Durand, Tom
Rahary, Adrien Ramanana
Ovsjanikov, Maks
contents Existing generative approaches for guided image synthesis of multi-object scenes typically rely on 2D controls in the image or text space. As a result, these methods struggle to maintain and respect consistent three-dimensional geometric structure, underlying the scene. In this paper, we propose a novel conditioning approach, training method and adapter network that can be plugged into pretrained text-to-image diffusion models. Our approach provides a way to endow such models with 3D-awareness, while leveraging their rich prior knowledge. Our method supports camera control, conditioning on explicit 3D geometries and, for the first time, accounts for the entire context of a scene, i.e., both on and off-screen items, to synthesize plausible and semantically rich images. Despite its multi-modal nature, our model is lightweight, requires a reasonable number of data for supervised learning and shows remarkable generalization power. We also introduce methods for intuitive and consistent image editing and restyling, e.g., by positioning, rotating or resizing individual objects in a scene. Our method integrates well within various image creation workflows and enables a richer set of applications compared to previous approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2507_03257
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LACONIC: A 3D Layout Adapter for Controllable Image Creation
Maillard, Léopold
Durand, Tom
Rahary, Adrien Ramanana
Ovsjanikov, Maks
Computer Vision and Pattern Recognition
Existing generative approaches for guided image synthesis of multi-object scenes typically rely on 2D controls in the image or text space. As a result, these methods struggle to maintain and respect consistent three-dimensional geometric structure, underlying the scene. In this paper, we propose a novel conditioning approach, training method and adapter network that can be plugged into pretrained text-to-image diffusion models. Our approach provides a way to endow such models with 3D-awareness, while leveraging their rich prior knowledge. Our method supports camera control, conditioning on explicit 3D geometries and, for the first time, accounts for the entire context of a scene, i.e., both on and off-screen items, to synthesize plausible and semantically rich images. Despite its multi-modal nature, our model is lightweight, requires a reasonable number of data for supervised learning and shows remarkable generalization power. We also introduce methods for intuitive and consistent image editing and restyling, e.g., by positioning, rotating or resizing individual objects in a scene. Our method integrates well within various image creation workflows and enables a richer set of applications compared to previous approaches.
title LACONIC: A 3D Layout Adapter for Controllable Image Creation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.03257