Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Chao, Wei, Tianyi, Yu, Nenghai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation
by: Zhang, Yabo, et al.
Published: (2026)
by: Zhang, Yabo, et al.
Published: (2026)
TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
by: Wei, Xinyu, et al.
Published: (2025)
by: Wei, Xinyu, et al.
Published: (2025)
Rethinking Multi-Condition DiTs: Eliminating Redundant Attention via Position-Alignment and Keyword-Scoping
by: Zhou, Chao, et al.
Published: (2026)
by: Zhou, Chao, et al.
Published: (2026)
Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
by: Pu, Yujiang, et al.
Published: (2025)
by: Pu, Yujiang, et al.
Published: (2025)
Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following
by: Feng, Yutong, et al.
Published: (2023)
by: Feng, Yutong, et al.
Published: (2023)
In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
by: Zhang, Zechuan, et al.
Published: (2025)
by: Zhang, Zechuan, et al.
Published: (2025)
Forge-and-Quench: Enhancing Image Generation for Higher Fidelity in Unified Multimodal Models
by: Zeng, Yanbing, et al.
Published: (2026)
by: Zeng, Yanbing, et al.
Published: (2026)
UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
by: Zhao, Haozhe, et al.
Published: (2024)
by: Zhao, Haozhe, et al.
Published: (2024)
Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
by: Feng, Kunyu, et al.
Published: (2025)
by: Feng, Kunyu, et al.
Published: (2025)
Enhanced Multi-Scale Cross-Attention for Person Image Generation
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Instruction-Free Tuning of Large Vision Language Models for Medical Instruction Following
by: Kang, Myeongkyun, et al.
Published: (2026)
by: Kang, Myeongkyun, et al.
Published: (2026)
Polaris: Scaling Up Instruction-Guided Image Generation Towards Millions of Personalized Style Needs
by: Chen, Zhi-Kai, et al.
Published: (2026)
by: Chen, Zhi-Kai, et al.
Published: (2026)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
by: Liu, Yexin, et al.
Published: (2025)
by: Liu, Yexin, et al.
Published: (2025)
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy
by: Yang, Te, et al.
Published: (2024)
by: Yang, Te, et al.
Published: (2024)
MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and Editing
by: Tian, Xueyun, et al.
Published: (2025)
by: Tian, Xueyun, et al.
Published: (2025)
Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional Videos
by: Nagasinghe, Kumaranage Ravindu Yasas, et al.
Published: (2024)
by: Nagasinghe, Kumaranage Ravindu Yasas, et al.
Published: (2024)
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning
by: Xu, Zhiyang, et al.
Published: (2024)
by: Xu, Zhiyang, et al.
Published: (2024)
DreamVE: Unified Instruction-based Image and Video Editing
by: Xia, Bin, et al.
Published: (2025)
by: Xia, Bin, et al.
Published: (2025)
PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation
by: Wang, Ao, et al.
Published: (2024)
by: Wang, Ao, et al.
Published: (2024)
Keypoint-Integrated Instruction-Following Data Generation for Enhanced Human Pose and Action Understanding in Multimodal Models
by: Zhang, Dewen, et al.
Published: (2024)
by: Zhang, Dewen, et al.
Published: (2024)
Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention Regulation in Diffusion Models
by: Zhang, Yang, et al.
Published: (2024)
by: Zhang, Yang, et al.
Published: (2024)
How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing
by: Zhang, Huanyu, et al.
Published: (2026)
by: Zhang, Huanyu, et al.
Published: (2026)
Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation
by: Wei, Tianyi, et al.
Published: (2024)
by: Wei, Tianyi, et al.
Published: (2024)
InsightEdit: Towards Better Instruction Following for Image Editing
by: Xu, Yingjing, et al.
Published: (2024)
by: Xu, Yingjing, et al.
Published: (2024)
IF-VidCap: Can Video Caption Models Follow Instructions?
by: Li, Shihao, et al.
Published: (2025)
by: Li, Shihao, et al.
Published: (2025)
Draw ALL Your Imagine: A Holistic Benchmark and Agent Framework for Complex Instruction-based Image Generation
by: Zhou, Yucheng, et al.
Published: (2025)
by: Zhou, Yucheng, et al.
Published: (2025)
Watch Your Steps: Local Image and Scene Editing by Text Instructions
by: Mirzaei, Ashkan, et al.
Published: (2023)
by: Mirzaei, Ashkan, et al.
Published: (2023)
Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model
by: Wang, Shijian, et al.
Published: (2024)
by: Wang, Shijian, et al.
Published: (2024)
Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
by: Chen, Harold Haodong, et al.
Published: (2025)
by: Chen, Harold Haodong, et al.
Published: (2025)
Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension Scoring
by: Xie, Qizhi, et al.
Published: (2025)
by: Xie, Qizhi, et al.
Published: (2025)
A High-Quality Text-Rich Image Instruction Tuning Dataset via Hybrid Instruction Generation
by: Zhou, Shijie, et al.
Published: (2024)
by: Zhou, Shijie, et al.
Published: (2024)
Learn Your Scales: Towards Scale-Consistent Generative Novel View Synthesis
by: Forghani, Fereshteh, et al.
Published: (2025)
by: Forghani, Fereshteh, et al.
Published: (2025)
UniBlendNet: Unified Global, Multi-Scale, and Region-Adaptive Modeling for Ambient Lighting Normalization
by: Dai, Jiatao, et al.
Published: (2026)
by: Dai, Jiatao, et al.
Published: (2026)
Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models
by: Huang, Yupan, et al.
Published: (2023)
by: Huang, Yupan, et al.
Published: (2023)
VII: Visual Instruction Injection for Jailbreaking Image-to-Video Generation Models
by: Zheng, Bowen, et al.
Published: (2026)
by: Zheng, Bowen, et al.
Published: (2026)
OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models
by: Dong, Xuanzhao, et al.
Published: (2026)
by: Dong, Xuanzhao, et al.
Published: (2026)
Efficient Inference of Vision Instruction-Following Models with Elastic Cache
by: Liu, Zuyan, et al.
Published: (2024)
by: Liu, Zuyan, et al.
Published: (2024)
DeCoT: Decomposing Complex Instructions for Enhanced Text-to-Image Generation with Large Language Models
by: Lin, Xiaochuan, et al.
Published: (2025)
by: Lin, Xiaochuan, et al.
Published: (2025)
CREval: An Automated Interpretable Evaluation for Creative Image Manipulation under Complex Instructions
by: Wang, Chonghuinan, et al.
Published: (2026)
by: Wang, Chonghuinan, et al.
Published: (2026)
Advancing Aesthetic Image Generation via Composition Transfer
by: Zou, Kai, et al.
Published: (2026)
by: Zou, Kai, et al.
Published: (2026)
Similar Items
-
Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation
by: Zhang, Yabo, et al.
Published: (2026) -
TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
by: Wei, Xinyu, et al.
Published: (2025) -
Rethinking Multi-Condition DiTs: Eliminating Redundant Attention via Position-Alignment and Keyword-Scoping
by: Zhou, Chao, et al.
Published: (2026) -
Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
by: Pu, Yujiang, et al.
Published: (2025) -
Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following
by: Feng, Yutong, et al.
Published: (2023)