Compositional Image Synthesis with Inference-Time Scaling

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ji, Minsuk, Lee, Sanghyeok, Ahn, Namhyuk
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912984661491712
author Ji, Minsuk
Lee, Sanghyeok
Ahn, Namhyuk
author_facet Ji, Minsuk
Lee, Sanghyeok
Ahn, Namhyuk
contents Despite their impressive realism, modern text-to-image models still struggle with compositionality, often failing to render accurate object counts, attributes, and spatial relations. To address this challenge, we present a training-free framework that combines an object-centric approach with self-refinement to improve layout faithfulness while preserving aesthetic quality. Specifically, we leverage large language models (LLMs) to synthesize explicit layouts from input prompts, and we inject these layouts into the image generation process, where a object-centric vision-language model (VLM) judge reranks multiple candidates to select the most prompt-aligned outcome iteratively. By unifying explicit layout-grounding with self-refine-based inference-time scaling, our framework achieves stronger scene alignment with prompts compared to recent text-to-image models. The code are available at https://github.com/gcl-inha/ReFocus.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24133
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Compositional Image Synthesis with Inference-Time Scaling
Ji, Minsuk
Lee, Sanghyeok
Ahn, Namhyuk
Computer Vision and Pattern Recognition
Artificial Intelligence
Despite their impressive realism, modern text-to-image models still struggle with compositionality, often failing to render accurate object counts, attributes, and spatial relations. To address this challenge, we present a training-free framework that combines an object-centric approach with self-refinement to improve layout faithfulness while preserving aesthetic quality. Specifically, we leverage large language models (LLMs) to synthesize explicit layouts from input prompts, and we inject these layouts into the image generation process, where a object-centric vision-language model (VLM) judge reranks multiple candidates to select the most prompt-aligned outcome iteratively. By unifying explicit layout-grounding with self-refine-based inference-time scaling, our framework achieves stronger scene alignment with prompts compared to recent text-to-image models. The code are available at https://github.com/gcl-inha/ReFocus.
title Compositional Image Synthesis with Inference-Time Scaling
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.24133