UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kang, Wonjun, Ahn, Byeongkeun, Lee, Minjae, Galim, Kevin, Oh, Seunghyuk, Koo, Hyung Il, Cho, Nam Ik
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913979138310144
author Kang, Wonjun
Ahn, Byeongkeun
Lee, Minjae
Galim, Kevin
Oh, Seunghyuk
Koo, Hyung Il
Cho, Nam Ik
author_facet Kang, Wonjun
Ahn, Byeongkeun
Lee, Minjae
Galim, Kevin
Oh, Seunghyuk
Koo, Hyung Il
Cho, Nam Ik
contents Text-to-image (T2I) generation has been actively studied using Diffusion Models and Autoregressive Models. Recently, Masked Generative Transformers have gained attention as an alternative to Autoregressive Models to overcome the inherent limitations of causal attention and autoregressive decoding through bidirectional attention and parallel decoding, enabling efficient and high-quality image generation. However, compositional T2I generation remains challenging, as even state-of-the-art Diffusion Models often fail to accurately bind attributes and achieve proper text-image alignment. While Diffusion Models have been extensively studied for this issue, Masked Generative Transformers exhibit similar limitations but have not been explored in this context. To address this, we propose Unmasking with Contrastive Attention Guidance (UNCAGE), a novel training-free method that improves compositional fidelity by leveraging attention maps to prioritize the unmasking of tokens that clearly represent individual objects. UNCAGE consistently improves performance in both quantitative and qualitative evaluations across multiple benchmarks and metrics, with negligible inference overhead. Our code is available at https://github.com/furiosa-ai/uncage.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05399
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
Kang, Wonjun
Ahn, Byeongkeun
Lee, Minjae
Galim, Kevin
Oh, Seunghyuk
Koo, Hyung Il
Cho, Nam Ik
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Text-to-image (T2I) generation has been actively studied using Diffusion Models and Autoregressive Models. Recently, Masked Generative Transformers have gained attention as an alternative to Autoregressive Models to overcome the inherent limitations of causal attention and autoregressive decoding through bidirectional attention and parallel decoding, enabling efficient and high-quality image generation. However, compositional T2I generation remains challenging, as even state-of-the-art Diffusion Models often fail to accurately bind attributes and achieve proper text-image alignment. While Diffusion Models have been extensively studied for this issue, Masked Generative Transformers exhibit similar limitations but have not been explored in this context. To address this, we propose Unmasking with Contrastive Attention Guidance (UNCAGE), a novel training-free method that improves compositional fidelity by leveraging attention maps to prioritize the unmasking of tokens that clearly represent individual objects. UNCAGE consistently improves performance in both quantitative and qualitative evaluations across multiple benchmarks and metrics, with negligible inference overhead. Our code is available at https://github.com/furiosa-ai/uncage.
title UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2508.05399