Controllable Image Generation with Composed Parallel Token Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stirling, Jamie, Al-Moubayed, Noura, Willcocks, Chris G., Shum, Hubert P. H.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915918140932096
author Stirling, Jamie
Al-Moubayed, Noura
Willcocks, Chris G.
Shum, Hubert P. H.
author_facet Stirling, Jamie
Al-Moubayed, Noura
Willcocks, Chris G.
Shum, Hubert P. H.
contents Conditional discrete generative models struggle to faithfully compose multiple input conditions. To address this, we derive a theoretically-grounded formulation for composing discrete probabilistic generative processes, with masked generation (absorbing diffusion) as a special case. Our formulation enables precise specification of novel combinations and numbers of input conditions that lie outside the training data, with concept weighting enabling emphasis or negation of individual conditions. In synergy with the richly compositional learned vocabulary of VQ-VAE and VQ-GAN, our method attains a $63.4\%$ relative reduction in error rate compared to the previous state-of-the-art, averaged across 3 datasets (positional CLEVR, relational CLEVR and FFHQ), simultaneously obtaining an average absolute FID improvement of $-9.58$. Meanwhile, our method offers a $2.3\times$ to $12\times$ real-time speed-up over comparable methods, and is readily applied to an open pre-trained discrete text-to-image model for fine-grained control of text-to-image generation.
format Preprint
id arxiv_https___arxiv_org_abs_2405_06535
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Controllable Image Generation with Composed Parallel Token Prediction
Stirling, Jamie
Al-Moubayed, Noura
Willcocks, Chris G.
Shum, Hubert P. H.
Computer Vision and Pattern Recognition
Machine Learning
Conditional discrete generative models struggle to faithfully compose multiple input conditions. To address this, we derive a theoretically-grounded formulation for composing discrete probabilistic generative processes, with masked generation (absorbing diffusion) as a special case. Our formulation enables precise specification of novel combinations and numbers of input conditions that lie outside the training data, with concept weighting enabling emphasis or negation of individual conditions. In synergy with the richly compositional learned vocabulary of VQ-VAE and VQ-GAN, our method attains a $63.4\%$ relative reduction in error rate compared to the previous state-of-the-art, averaged across 3 datasets (positional CLEVR, relational CLEVR and FFHQ), simultaneously obtaining an average absolute FID improvement of $-9.58$. Meanwhile, our method offers a $2.3\times$ to $12\times$ real-time speed-up over comparable methods, and is readily applied to an open pre-trained discrete text-to-image model for fine-grained control of text-to-image generation.
title Controllable Image Generation with Composed Parallel Token Prediction
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2405.06535