VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Shaojin, Ding, Fei, Huang, Mengqi, Liu, Wei, He, Qian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Less-to-More Generalization: Unlocking More Controllability by In-Context Generation
by: Wu, Shaojin, et al.
Published: (2025)
by: Wu, Shaojin, et al.
Published: (2025)
UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward
by: Cheng, Yufeng, et al.
Published: (2025)
by: Cheng, Yufeng, et al.
Published: (2025)
USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning
by: Wu, Shaojin, et al.
Published: (2025)
by: Wu, Shaojin, et al.
Published: (2025)
AI-T2I: Aggregating-and-Isolating Cross-Attention to Diffusion Models for Text-to-Image Synthesis
by: Cao, Shipeng, et al.
Published: (2026)
by: Cao, Shipeng, et al.
Published: (2026)
RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image Customization
by: Huang, Mengqi, et al.
Published: (2024)
by: Huang, Mengqi, et al.
Published: (2024)
RealCustom++: Representing Images as Real Textual Word for Real-Time Customization
by: Mao, Zhendong, et al.
Published: (2024)
by: Mao, Zhendong, et al.
Published: (2024)
Local Conditional Controlling for Text-to-Image Diffusion Models
by: Zhao, Yibo, et al.
Published: (2023)
by: Zhao, Yibo, et al.
Published: (2023)
Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing
by: Liu, Bingyan, et al.
Published: (2024)
by: Liu, Bingyan, et al.
Published: (2024)
Stream-T1: Test-Time Scaling for Streaming Video Generation
by: Tu, Yijing, et al.
Published: (2026)
by: Tu, Yijing, et al.
Published: (2026)
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
by: Wu, Bin, et al.
Published: (2026)
by: Wu, Bin, et al.
Published: (2026)
AID: Attention Interpolation of Text-to-Image Diffusion
by: He, Qiyuan, et al.
Published: (2024)
by: He, Qiyuan, et al.
Published: (2024)
Contrastive Prompts Improve Disentanglement in Text-to-Image Diffusion Models
by: Wu, Chen, et al.
Published: (2024)
by: Wu, Chen, et al.
Published: (2024)
Scaling Down Text Encoders of Text-to-Image Diffusion Models
by: Wang, Lifu, et al.
Published: (2025)
by: Wang, Lifu, et al.
Published: (2025)
Lance: Unified Multimodal Modeling by Multi-Task Synergy
by: Fu, Fengyi, et al.
Published: (2026)
by: Fu, Fengyi, et al.
Published: (2026)
MIST: Mitigating Intersectional Bias with Disentangled Cross-Attention Editing in Text-to-Image Diffusion Models
by: Yesiltepe, Hidir, et al.
Published: (2024)
by: Yesiltepe, Hidir, et al.
Published: (2024)
NativeTok: Native Visual Tokenization for Improved Image Generation
by: Wu, Bin, et al.
Published: (2026)
by: Wu, Bin, et al.
Published: (2026)
Text-guided Controllable Diffusion for Realistic Camouflage Images Generation
by: Qian, Yuhang, et al.
Published: (2025)
by: Qian, Yuhang, et al.
Published: (2025)
Layered Rendering Diffusion Model for Controllable Zero-Shot Image Synthesis
by: Qi, Zipeng, et al.
Published: (2023)
by: Qi, Zipeng, et al.
Published: (2023)
LayerDiffusion: Layered Controlled Image Editing with Diffusion Models
by: Li, Pengzhi, et al.
Published: (2023)
by: Li, Pengzhi, et al.
Published: (2023)
Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention Regulation in Diffusion Models
by: Zhang, Yang, et al.
Published: (2024)
by: Zhang, Yang, et al.
Published: (2024)
ECNet: Effective Controllable Text-to-Image Diffusion Models
by: Li, Sicheng, et al.
Published: (2024)
by: Li, Sicheng, et al.
Published: (2024)
Object-Conditioned Energy-Based Attention Map Alignment in Text-to-Image Diffusion Models
by: Zhang, Yasi, et al.
Published: (2024)
by: Zhang, Yasi, et al.
Published: (2024)
Bokeh Diffusion: Defocus Blur Control in Text-to-Image Diffusion Models
by: Fortes, Armando, et al.
Published: (2025)
by: Fortes, Armando, et al.
Published: (2025)
From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models
by: Xiao, Changming, et al.
Published: (2023)
by: Xiao, Changming, et al.
Published: (2023)
Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image Translation
by: Gao, Xiang, et al.
Published: (2024)
by: Gao, Xiang, et al.
Published: (2024)
CTCal: Rethinking Text-to-Image Diffusion Models via Cross-Timestep Self-Calibration
by: Guo, Xiefan, et al.
Published: (2026)
by: Guo, Xiefan, et al.
Published: (2026)
Predicated Diffusion: Predicate Logic-Based Attention Guidance for Text-to-Image Diffusion Models
by: Sueyoshi, Kota, et al.
Published: (2023)
by: Sueyoshi, Kota, et al.
Published: (2023)
Dynamic Attention Analysis for Backdoor Detection in Text-to-Image Diffusion Models
by: Wang, Zhongqi, et al.
Published: (2025)
by: Wang, Zhongqi, et al.
Published: (2025)
Analyzing and Improving Fast Sampling of Text-to-Image Diffusion Models
by: Zhou, Zhenyu, et al.
Published: (2026)
by: Zhou, Zhenyu, et al.
Published: (2026)
UniCtrl: Improving the Spatiotemporal Consistency of Text-to-Video Diffusion Models via Training-Free Unified Attention Control
by: Xia, Tian, et al.
Published: (2024)
by: Xia, Tian, et al.
Published: (2024)
Image Augmentation with Controlled Diffusion for Weakly-Supervised Semantic Segmentation
by: Wu, Wangyu, et al.
Published: (2023)
by: Wu, Wangyu, et al.
Published: (2023)
Debiasing Text-to-Image Diffusion Models
by: He, Ruifei, et al.
Published: (2024)
by: He, Ruifei, et al.
Published: (2024)
CIV-DG: Conditional Instrumental Variables for Domain Generalization in Medical Imaging
by: Bai, Shaojin, et al.
Published: (2026)
by: Bai, Shaojin, et al.
Published: (2026)
Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative Grounding
by: Li, Hongyu, et al.
Published: (2024)
by: Li, Hongyu, et al.
Published: (2024)
Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation
by: Wang, Wenjing, et al.
Published: (2023)
by: Wang, Wenjing, et al.
Published: (2023)
HDGlyph: A Hierarchical Disentangled Glyph-Based Framework for Long-Tail Text Rendering in Diffusion Models
by: Zhuang, Shuhan, et al.
Published: (2025)
by: Zhuang, Shuhan, et al.
Published: (2025)
Improving Long-Text Alignment for Text-to-Image Diffusion Models
by: Liu, Luping, et al.
Published: (2024)
by: Liu, Luping, et al.
Published: (2024)
DreamO: A Unified Framework for Image Customization
by: Mou, Chong, et al.
Published: (2025)
by: Mou, Chong, et al.
Published: (2025)
Controllable Generation with Text-to-Image Diffusion Models: A Survey
by: Cao, Pu, et al.
Published: (2024)
by: Cao, Pu, et al.
Published: (2024)
Immunizing Images from Text to Image Editing via Adversarial Cross-Attention
by: Trippodo, Matteo, et al.
Published: (2025)
by: Trippodo, Matteo, et al.
Published: (2025)
Similar Items
-
Less-to-More Generalization: Unlocking More Controllability by In-Context Generation
by: Wu, Shaojin, et al.
Published: (2025) -
UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward
by: Cheng, Yufeng, et al.
Published: (2025) -
USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning
by: Wu, Shaojin, et al.
Published: (2025) -
AI-T2I: Aggregating-and-Isolating Cross-Attention to Diffusion Models for Text-to-Image Synthesis
by: Cao, Shipeng, et al.
Published: (2026) -
RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image Customization
by: Huang, Mengqi, et al.
Published: (2024)