ConsistCompose: Unified Multimodal Layout Control for Image Composition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Xuanke, Li, Boxuan, Han, Xiaoyang, Cai, Zhongang, Yang, Lei, Wang, Quan, Lin, Dahua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910053420761088
author Shi, Xuanke
Li, Boxuan
Han, Xiaoyang
Cai, Zhongang
Yang, Lei
Wang, Quan
Lin, Dahua
author_facet Shi, Xuanke
Li, Boxuan
Han, Xiaoyang
Cai, Zhongang
Yang, Lei
Wang, Quan
Lin, Dahua
contents Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding-aligning language with image regions-while their generative counterpart, linguistic-embedded layout-grounded generation (LELG) for layout-controllable multi-instance generation, remains underexplored and limits precise compositional control. We present ConsistCompose, a unified multimodal framework that embeds layout coordinates directly into language prompts, enabling layout-controlled multi-instance image generation from Interleaved Image-Text within a single generative interface. We further construct ConsistCompose3M, a 3.4M multi-instance generation dataset with layout and identity annotations (2.6M text-guided and 0.8M image-guided data pairs) that provides large-scale supervision for layout-conditioned generation. Within this framework, LELG is instantiated through instance-coordinate binding prompts and coordinate-aware classifier-free guidance, which translate linguistic layout cues into precise spatial control without task-specific branches. Experiments on COCO-Position and MS-Bench show that ConsistCompose substantially improves spatial accuracy over layout-controlled baselines while preserving identity fidelity and competitive general multimodal understanding, establishing a unified paradigm for layout-controllable multimodal image generation.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18333
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ConsistCompose: Unified Multimodal Layout Control for Image Composition
Shi, Xuanke
Li, Boxuan
Han, Xiaoyang
Cai, Zhongang
Yang, Lei
Wang, Quan
Lin, Dahua
Computer Vision and Pattern Recognition
Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding-aligning language with image regions-while their generative counterpart, linguistic-embedded layout-grounded generation (LELG) for layout-controllable multi-instance generation, remains underexplored and limits precise compositional control. We present ConsistCompose, a unified multimodal framework that embeds layout coordinates directly into language prompts, enabling layout-controlled multi-instance image generation from Interleaved Image-Text within a single generative interface. We further construct ConsistCompose3M, a 3.4M multi-instance generation dataset with layout and identity annotations (2.6M text-guided and 0.8M image-guided data pairs) that provides large-scale supervision for layout-conditioned generation. Within this framework, LELG is instantiated through instance-coordinate binding prompts and coordinate-aware classifier-free guidance, which translate linguistic layout cues into precise spatial control without task-specific branches. Experiments on COCO-Position and MS-Bench show that ConsistCompose substantially improves spatial accuracy over layout-controlled baselines while preserving identity fidelity and competitive general multimodal understanding, establishing a unified paradigm for layout-controllable multimodal image generation.
title ConsistCompose: Unified Multimodal Layout Control for Image Composition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.18333