Chain-of-Image Generation: Toward Monitorable and Controllable Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Young Kyung, Schlesinger, Oded, Zhao, Yuzhou, Di Martino, J. Matias, Sapiro, Guillermo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914190074052608
author Kim, Young Kyung
Schlesinger, Oded
Zhao, Yuzhou
Di Martino, J. Matias
Sapiro, Guillermo
author_facet Kim, Young Kyung
Schlesinger, Oded
Zhao, Yuzhou
Di Martino, J. Matias
Sapiro, Guillermo
contents While state-of-the-art image generation models achieve remarkable visual quality, their internal generative processes remain a "black box." This opacity limits human observation and intervention, and poses a barrier to ensuring model reliability, safety, and control. Furthermore, their non-human-like workflows make them difficult for human observers to interpret. To address this, we introduce the Chain-of-Image Generation (CoIG) framework, which reframes image generation as a sequential, semantic process analogous to how humans create art. Similar to the advantages in monitorability and performance that Chain-of-Thought (CoT) brought to large language models (LLMs), CoIG can produce equivalent benefits in text-to-image generation. CoIG utilizes an LLM to decompose a complex prompt into a sequence of simple, step-by-step instructions. The image generation model then executes this plan by progressively generating and editing the image. Each step focuses on a single semantic entity, enabling direct monitoring. We formally assess this property using two novel metrics: CoIG Readability, which evaluates the clarity of each intermediate step via its corresponding output; and Causal Relevance, which quantifies the impact of each procedural step on the final generated image. We further show that our framework mitigates entity collapse by decomposing the complex generation task into simple subproblems, analogous to the procedural reasoning employed by CoT. Our experimental results indicate that CoIG substantially enhances quantitative monitorability while achieving competitive compositional robustness compared to established baseline models. The framework is model-agnostic and can be integrated with any image generation model.
format Preprint
id arxiv_https___arxiv_org_abs_2512_08645
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Chain-of-Image Generation: Toward Monitorable and Controllable Image Generation
Kim, Young Kyung
Schlesinger, Oded
Zhao, Yuzhou
Di Martino, J. Matias
Sapiro, Guillermo
Computer Vision and Pattern Recognition
While state-of-the-art image generation models achieve remarkable visual quality, their internal generative processes remain a "black box." This opacity limits human observation and intervention, and poses a barrier to ensuring model reliability, safety, and control. Furthermore, their non-human-like workflows make them difficult for human observers to interpret. To address this, we introduce the Chain-of-Image Generation (CoIG) framework, which reframes image generation as a sequential, semantic process analogous to how humans create art. Similar to the advantages in monitorability and performance that Chain-of-Thought (CoT) brought to large language models (LLMs), CoIG can produce equivalent benefits in text-to-image generation. CoIG utilizes an LLM to decompose a complex prompt into a sequence of simple, step-by-step instructions. The image generation model then executes this plan by progressively generating and editing the image. Each step focuses on a single semantic entity, enabling direct monitoring. We formally assess this property using two novel metrics: CoIG Readability, which evaluates the clarity of each intermediate step via its corresponding output; and Causal Relevance, which quantifies the impact of each procedural step on the final generated image. We further show that our framework mitigates entity collapse by decomposing the complex generation task into simple subproblems, analogous to the procedural reasoning employed by CoT. Our experimental results indicate that CoIG substantially enhances quantitative monitorability while achieving competitive compositional robustness compared to established baseline models. The framework is model-agnostic and can be integrated with any image generation model.
title Chain-of-Image Generation: Toward Monitorable and Controllable Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.08645