UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Wonjun, Ahn, Byeongkeun, Lee, Minjae, Galim, Kevin, Oh, Seunghyuk, Koo, Hyung Il, Cho, Nam Ik |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Counting Guidance for High Fidelity Text-to-Image Synthesis
by: Kang, Wonjun, et al.
Published: (2023)
by: Kang, Wonjun, et al.
Published: (2023)
TABED: Test-Time Adaptive Ensemble Drafting for Robust Speculative Decoding in LVLMs
by: Lee, Minjae, et al.
Published: (2026)
by: Lee, Minjae, et al.
Published: (2026)
State-offset Tuning: State-based Parameter-Efficient Fine-Tuning for State Space Models
by: Kang, Wonjun, et al.
Published: (2025)
by: Kang, Wonjun, et al.
Published: (2025)
Eta Inversion: Designing an Optimal Eta Function for Diffusion-based Real Image Editing
by: Kang, Wonjun, et al.
Published: (2024)
by: Kang, Wonjun, et al.
Published: (2024)
ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs
by: Kang, Wonjun, et al.
Published: (2025)
by: Kang, Wonjun, et al.
Published: (2025)
Draft-based Approximate Inference for LLMs
by: Galim, Kevin, et al.
Published: (2025)
by: Galim, Kevin, et al.
Published: (2025)
Parameter-Efficient Fine-Tuning of State Space Models
by: Galim, Kevin, et al.
Published: (2024)
by: Galim, Kevin, et al.
Published: (2024)
Can MLLMs Perform Text-to-Image In-Context Learning?
by: Zeng, Yuchen, et al.
Published: (2024)
by: Zeng, Yuchen, et al.
Published: (2024)
TM-BSN: Triangular-Masked Blind-Spot Network for Real-World Self-Supervised Image Denoising
by: Park, Junyoung, et al.
Published: (2026)
by: Park, Junyoung, et al.
Published: (2026)
Efficient Attention-Sharing Information Distillation Transformer for Lightweight Single Image Super-Resolution
by: Park, Karam, et al.
Published: (2025)
by: Park, Karam, et al.
Published: (2025)
Semantic Watermarking Reinvented: Enhancing Robustness and Generation Quality with Fourier Integrity
by: Lee, Sung Ju, et al.
Published: (2025)
by: Lee, Sung Ju, et al.
Published: (2025)
Towards Controllable Real Image Denoising with Camera Parameters
by: Oh, Youngjin, et al.
Published: (2025)
by: Oh, Youngjin, et al.
Published: (2025)
Content-Aware Preserving Image Generation
by: Le, Giang H., et al.
Published: (2024)
by: Le, Giang H., et al.
Published: (2024)
Generalized Zero-Shot Learning for Point Cloud Segmentation with Evidence-Based Dynamic Calibration
by: Kim, Hyeonseok, et al.
Published: (2025)
by: Kim, Hyeonseok, et al.
Published: (2025)
Generalized Class Discovery in Instance Segmentation
by: Hoang, Cuong Manh, et al.
Published: (2025)
by: Hoang, Cuong Manh, et al.
Published: (2025)
Unsupervised Contrastive Learning Using Out-Of-Distribution Data for Long-Tailed Dataset
by: Hoang, Cuong Manh, et al.
Published: (2025)
by: Hoang, Cuong Manh, et al.
Published: (2025)
PhaseMark: A Post-hoc, Optimization-Free Watermarking of AI-generated Images in the Latent Frequency Domain
by: Lee, Sung Ju, et al.
Published: (2026)
by: Lee, Sung Ju, et al.
Published: (2026)
Transformers in the Dark: Navigating Unknown Search Spaces via Bandit Feedback
by: Kim, Jungtaek, et al.
Published: (2026)
by: Kim, Jungtaek, et al.
Published: (2026)
DINOLight: Robust Ambient Light Normalization with Self-supervised Visual Prior Integration
by: Oh, Youngjin, et al.
Published: (2026)
by: Oh, Youngjin, et al.
Published: (2026)
Enhancing Multi-Exposure High Dynamic Range Imaging with Overlapped Codebook for Improved Representation Learning
by: Lee, Keuntek, et al.
Published: (2025)
by: Lee, Keuntek, et al.
Published: (2025)
Lightweight and Fast Real-time Image Enhancement via Decomposition of the Spatial-aware Lookup Tables
by: Kim, Wontae, et al.
Published: (2025)
by: Kim, Wontae, et al.
Published: (2025)
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
by: Ko, Jungmin, et al.
Published: (2026)
by: Ko, Jungmin, et al.
Published: (2026)
Optimized Feature Generation for Tabular Data via LLMs with Decision Tree Reasoning
by: Nam, Jaehyun, et al.
Published: (2024)
by: Nam, Jaehyun, et al.
Published: (2024)
Improving Weakly-Supervised Object Localization Using Adversarial Erasing and Pseudo Label
by: Kang, Byeongkeun, et al.
Published: (2024)
by: Kang, Byeongkeun, et al.
Published: (2024)
High Fidelity Text to Image Generation with Contrastive Alignment and Structural Guidance
by: Gao, Danyi
Published: (2025)
by: Gao, Danyi
Published: (2025)
Tidiness Score-Guided Monte Carlo Tree Search for Visual Tabletop Rearrangement
by: Kee, Hogun, et al.
Published: (2025)
by: Kee, Hogun, et al.
Published: (2025)
CountCluster: Training-Free Object Quantity Guidance with Cross-Attention Map Clustering for Text-to-Image Generation
by: Lee, Joohyeon, et al.
Published: (2025)
by: Lee, Joohyeon, et al.
Published: (2025)
DarkVRAI: Capture-Condition Conditioning and Burst-Order Selective Scan for Low-light RAW Video Denoising
by: Oh, Youngjin, et al.
Published: (2025)
by: Oh, Youngjin, et al.
Published: (2025)
OmniLight: One Model to Rule All Lighting Conditions
by: Oh, Youngjin, et al.
Published: (2026)
by: Oh, Youngjin, et al.
Published: (2026)
Pseudochaotic Many-Body Dynamics as a Pseudorandom State Generator
by: Lee, Wonjun, et al.
Published: (2024)
by: Lee, Wonjun, et al.
Published: (2024)
Completely Weakly Supervised Class-Incremental Learning for Semantic Segmentation
by: Kim, David Minkwan, et al.
Published: (2025)
by: Kim, David Minkwan, et al.
Published: (2025)
Enhancing Long-Term Person Re-Identification Using Global, Local Body Part, and Head Streams
by: Thanh, Duy Tran, et al.
Published: (2024)
by: Thanh, Duy Tran, et al.
Published: (2024)
Stop Jostling: Adaptive Negative Sampling Reduces the Marginalization of Low-Resource Language Tokens by Cross-Entropy Loss
by: Turumtaev, Galim
Published: (2026)
by: Turumtaev, Galim
Published: (2026)
DAFT-GAN: Dual Affine Transformation Generative Adversarial Network for Text-Guided Image Inpainting
by: Lee, Jihoon, et al.
Published: (2024)
by: Lee, Jihoon, et al.
Published: (2024)
Masked Generative Story Transformer with Character Guidance and Caption Augmentation
by: Papadimitriou, Christos, et al.
Published: (2024)
by: Papadimitriou, Christos, et al.
Published: (2024)
Training-Free Safe Text Embedding Guidance for Text-to-Image Diffusion Models
by: Na, Byeonghu, et al.
Published: (2025)
by: Na, Byeonghu, et al.
Published: (2025)
Prompt-Based Safety Guidance Is Ineffective for Unlearned Text-to-Image Diffusion Models
by: Shin, Jiwoo, et al.
Published: (2025)
by: Shin, Jiwoo, et al.
Published: (2025)
Semantic Guidance Tuning for Text-To-Image Diffusion Models
by: Kang, Hyun, et al.
Published: (2023)
by: Kang, Hyun, et al.
Published: (2023)
Unlocking the Capabilities of Masked Generative Models for Image Synthesis via Self-Guidance
by: Hur, Jiwan, et al.
Published: (2024)
by: Hur, Jiwan, et al.
Published: (2024)
MSTA3D: Multi-scale Twin-attention for 3D Instance Segmentation
by: Tran, Duc Dang Trung, et al.
Published: (2024)
by: Tran, Duc Dang Trung, et al.
Published: (2024)
Similar Items
-
Counting Guidance for High Fidelity Text-to-Image Synthesis
by: Kang, Wonjun, et al.
Published: (2023) -
TABED: Test-Time Adaptive Ensemble Drafting for Robust Speculative Decoding in LVLMs
by: Lee, Minjae, et al.
Published: (2026) -
State-offset Tuning: State-based Parameter-Efficient Fine-Tuning for State Space Models
by: Kang, Wonjun, et al.
Published: (2025) -
Eta Inversion: Designing an Optimal Eta Function for Diffusion-based Real Image Editing
by: Kang, Wonjun, et al.
Published: (2024) -
ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMs
by: Kang, Wonjun, et al.
Published: (2025)