Controllable Coupled Image Generation via Diffusion Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yuan, Chenfei, Jia, Nanshan, Li, Hangqi, Glynn, Peter W., Zheng, Zeyu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913884764372992
author Yuan, Chenfei
Jia, Nanshan
Li, Hangqi
Glynn, Peter W.
Zheng, Zeyu
author_facet Yuan, Chenfei
Jia, Nanshan
Li, Hangqi
Glynn, Peter W.
Zheng, Zeyu
contents We provide an attention-level control method for the task of coupled image generation, where "coupled" means that multiple simultaneously generated images are expected to have the same or very similar backgrounds. While backgrounds coupled, the centered objects in the generated images are still expected to enjoy the flexibility raised from different text prompts. The proposed method disentangles the background and entity components in the model's cross-attention modules, attached with a sequence of time-varying weight control parameters depending on the time step of sampling. We optimize this sequence of weight control parameters with a combined objective that assesses how coupled the backgrounds are as well as text-to-image alignment and overall visual quality. Empirical results demonstrate that our method outperforms existing approaches across these criteria.
format Preprint
id arxiv_https___arxiv_org_abs_2506_06826
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Controllable Coupled Image Generation via Diffusion Models
Yuan, Chenfei
Jia, Nanshan
Li, Hangqi
Glynn, Peter W.
Zheng, Zeyu
Computer Vision and Pattern Recognition
Artificial Intelligence
We provide an attention-level control method for the task of coupled image generation, where "coupled" means that multiple simultaneously generated images are expected to have the same or very similar backgrounds. While backgrounds coupled, the centered objects in the generated images are still expected to enjoy the flexibility raised from different text prompts. The proposed method disentangles the background and entity components in the model's cross-attention modules, attached with a sequence of time-varying weight control parameters depending on the time step of sampling. We optimize this sequence of weight control parameters with a combined objective that assesses how coupled the backgrounds are as well as text-to-image alignment and overall visual quality. Empirical results demonstrate that our method outperforms existing approaches across these criteria.
title Controllable Coupled Image Generation via Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.06826