MaskAttn-SDXL: Controllable Region-Level Text-To-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Chang, Yu, Chen, Jiahao, Cheng, Anzhe, Bogdan, Paul |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MaskAttn-UNet: A Mask Attention-Driven Framework for Universal Low-Resolution Image Segmentation
by: Cheng, Anzhe, et al.
Published: (2025)
by: Cheng, Anzhe, et al.
Published: (2025)
Eigen Neural Network: Unlocking Generalizable Vision with Eigenbasis
by: Cheng, Anzhe, et al.
Published: (2025)
by: Cheng, Anzhe, et al.
Published: (2025)
EMoE: Eigenbasis-Guided Routing for Mixture-of-Experts
by: Cheng, Anzhe, et al.
Published: (2026)
by: Cheng, Anzhe, et al.
Published: (2026)
SDXL-Lightning: Progressive Adversarial Diffusion Distillation
by: Lin, Shanchuan, et al.
Published: (2024)
by: Lin, Shanchuan, et al.
Published: (2024)
Improvements to SDXL in NovelAI Diffusion V3
by: Ossa, Juan, et al.
Published: (2024)
by: Ossa, Juan, et al.
Published: (2024)
MaskBit: Embedding-free Image Generation via Bit Tokens
by: Weber, Mark, et al.
Published: (2024)
by: Weber, Mark, et al.
Published: (2024)
Skews in the Phenomenon Space Hinder Generalization in Text-to-Image Generation
by: Chang, Yingshan, et al.
Published: (2024)
by: Chang, Yingshan, et al.
Published: (2024)
Progressive Compositionality in Text-to-Image Generative Models
by: Han, Evans Xu, et al.
Published: (2024)
by: Han, Evans Xu, et al.
Published: (2024)
UNCAGE: Contrastive Attention Guidance for Masked Generative Transformers in Text-to-Image Generation
by: Kang, Wonjun, et al.
Published: (2025)
by: Kang, Wonjun, et al.
Published: (2025)
Effective and Efficient Masked Image Generation Models
by: You, Zebin, et al.
Published: (2025)
by: You, Zebin, et al.
Published: (2025)
MaskVD: Region Masking for Efficient Video Object Detection
by: Sarkar, Sreetama, et al.
Published: (2024)
by: Sarkar, Sreetama, et al.
Published: (2024)
Cocktail: Mixing Multi-Modality Controls for Text-Conditional Image Generation
by: Hu, Minghui, et al.
Published: (2023)
by: Hu, Minghui, et al.
Published: (2023)
Structural Complexity of Brain MRI reveals age-associated patterns
by: Cheng, Anzhe, et al.
Published: (2026)
by: Cheng, Anzhe, et al.
Published: (2026)
MedSegFactory: Text-Guided Generation of Medical Image-Mask Pairs
by: Mao, Jiawei, et al.
Published: (2025)
by: Mao, Jiawei, et al.
Published: (2025)
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
by: Li, Chao, et al.
Published: (2026)
by: Li, Chao, et al.
Published: (2026)
AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image Generation
by: Pang, Lianyu, et al.
Published: (2024)
by: Pang, Lianyu, et al.
Published: (2024)
Disentangling Regional Primitives for Image Generation
by: Chen, Zhengting, et al.
Published: (2024)
by: Chen, Zhengting, et al.
Published: (2024)
Masked Generative Transformer Is What You Need for Image Editing
by: Chow, Wei, et al.
Published: (2026)
by: Chow, Wei, et al.
Published: (2026)
Ctrl-X: Controlling Structure and Appearance for Text-To-Image Generation Without Guidance
by: Lin, Kuan Heng, et al.
Published: (2024)
by: Lin, Kuan Heng, et al.
Published: (2024)
Downstream Task Guided Masking Learning in Masked Autoencoders Using Multi-Level Optimization
by: Guo, Han, et al.
Published: (2024)
by: Guo, Han, et al.
Published: (2024)
Model-Agnostic Gender Bias Control for Text-to-Image Generation via Sparse Autoencoder
by: Wu, Chao, et al.
Published: (2025)
by: Wu, Chao, et al.
Published: (2025)
EPIC: Efficient Predicate-Guided Inference-Time Control for Compositional Text-to-Image Generation
by: Mun, Sunung, et al.
Published: (2026)
by: Mun, Sunung, et al.
Published: (2026)
Cost-Aware Routing for Efficient Text-To-Image Generation
by: Li, Qinchan, et al.
Published: (2025)
by: Li, Qinchan, et al.
Published: (2025)
Multi-Level Feature Distillation of Joint Teachers Trained on Distinct Image Datasets
by: Iordache, Adrian, et al.
Published: (2024)
by: Iordache, Adrian, et al.
Published: (2024)
Exploring the Coordination of Frequency and Attention in Masked Image Modeling
by: Gui, Jie, et al.
Published: (2022)
by: Gui, Jie, et al.
Published: (2022)
Keypoint Aware Masked Image Modelling
by: Krishna, Madhava, et al.
Published: (2024)
by: Krishna, Madhava, et al.
Published: (2024)
Aligned but Stereotypical? The Hidden Influence of System Prompts on Social Bias in LVLM-Based Text-to-Image Models
by: Park, NaHyeon, et al.
Published: (2025)
by: Park, NaHyeon, et al.
Published: (2025)
Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation
by: Yariv, Guy, et al.
Published: (2025)
by: Yariv, Guy, et al.
Published: (2025)
Reward Incremental Learning in Text-to-Image Generation
by: Wang, Maorong, et al.
Published: (2024)
by: Wang, Maorong, et al.
Published: (2024)
Uncovering Regional Defaults from Photorealistic Forests in Text-to-Image Generation with DALL-E 2
by: Liu, Zilong, et al.
Published: (2024)
by: Liu, Zilong, et al.
Published: (2024)
Hawk: Leveraging Spatial Context for Faster Autoregressive Text-to-Image Generation
by: Chen, Zhi-Kai, et al.
Published: (2025)
by: Chen, Zhi-Kai, et al.
Published: (2025)
DICE: Discrete Inversion Enabling Controllable Editing for Multinomial Diffusion and Masked Generative Models
by: He, Xiaoxiao, et al.
Published: (2024)
by: He, Xiaoxiao, et al.
Published: (2024)
Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models
by: Choi, Hyesong, et al.
Published: (2024)
by: Choi, Hyesong, et al.
Published: (2024)
Expressive Text-to-Image Generation with Rich Text
by: Ge, Songwei, et al.
Published: (2023)
by: Ge, Songwei, et al.
Published: (2023)
Text-Driven Image Editing via Learnable Regions
by: Lin, Yuanze, et al.
Published: (2023)
by: Lin, Yuanze, et al.
Published: (2023)
Contrastive Masked Autoencoders for Character-Level Open-Set Writer Identification
by: Jiang, Xiaowei, et al.
Published: (2025)
by: Jiang, Xiaowei, et al.
Published: (2025)
Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps
by: Kim, Jeeyung, et al.
Published: (2024)
by: Kim, Jeeyung, et al.
Published: (2024)
MaskMedPaint: Masked Medical Image Inpainting with Diffusion Models for Mitigation of Spurious Correlations
by: Jin, Qixuan, et al.
Published: (2024)
by: Jin, Qixuan, et al.
Published: (2024)
What Shape Is Optimal for Masks in Text Removal?
by: Nakada, Hyakka, et al.
Published: (2025)
by: Nakada, Hyakka, et al.
Published: (2025)
Directional Textual Inversion for Personalized Text-to-Image Generation
by: Kim, Kunhee, et al.
Published: (2025)
by: Kim, Kunhee, et al.
Published: (2025)
Similar Items
-
MaskAttn-UNet: A Mask Attention-Driven Framework for Universal Low-Resolution Image Segmentation
by: Cheng, Anzhe, et al.
Published: (2025) -
Eigen Neural Network: Unlocking Generalizable Vision with Eigenbasis
by: Cheng, Anzhe, et al.
Published: (2025) -
EMoE: Eigenbasis-Guided Routing for Mixture-of-Experts
by: Cheng, Anzhe, et al.
Published: (2026) -
SDXL-Lightning: Progressive Adversarial Diffusion Distillation
by: Lin, Shanchuan, et al.
Published: (2024) -
Improvements to SDXL in NovelAI Diffusion V3
by: Ossa, Juan, et al.
Published: (2024)