Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Yutong, Gong, Biao, Chen, Di, Shen, Yujun, Liu, Yu, Zhou, Jingren |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mimir: Improving Video Diffusion Models for Precise Text Understanding
by: Tan, Shuai, et al.
Published: (2024)
by: Tan, Shuai, et al.
Published: (2024)
Learning Disentangled Identifiers for Action-Customized Text-to-Image Generation
by: Huang, Siteng, et al.
Published: (2023)
by: Huang, Siteng, et al.
Published: (2023)
Check, Locate, Rectify: A Training-Free Layout Calibration System for Text-to-Image Generation
by: Gong, Biao, et al.
Published: (2023)
by: Gong, Biao, et al.
Published: (2023)
ACE: All-round Creator and Editor Following Instructions via Diffusion Transformer
by: Han, Zhen, et al.
Published: (2024)
by: Han, Zhen, et al.
Published: (2024)
Zero-shot Image Editing with Reference Imitation
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
CoDance: An Unbind-Rebind Paradigm for Robust Multi-Subject Animation
by: Tan, Shuai, et al.
Published: (2026)
by: Tan, Shuai, et al.
Published: (2026)
Taming Stable Diffusion for Text to 360° Panorama Image Generation
by: Zhang, Cheng, et al.
Published: (2024)
by: Zhang, Cheng, et al.
Published: (2024)
Lipschitz Singularities in Diffusion Models
by: Yang, Zhantao, et al.
Published: (2023)
by: Yang, Zhantao, et al.
Published: (2023)
Towards More Accurate Diffusion Model Acceleration with A Timestep Tuner
by: Xia, Mengfei, et al.
Published: (2023)
by: Xia, Mengfei, et al.
Published: (2023)
UKnow: A Unified Knowledge Protocol with Multimodal Knowledge Graph Datasets for Reasoning and Vision-Language Pre-Training
by: Gong, Biao, et al.
Published: (2023)
by: Gong, Biao, et al.
Published: (2023)
Group Diffusion Transformers are Unsupervised Multitask Learners
by: Huang, Lianghua, et al.
Published: (2024)
by: Huang, Lianghua, et al.
Published: (2024)
TIP-Editor: An Accurate 3D Editor Following Both Text-Prompts And Image-Prompts
by: Zhuang, Jingyu, et al.
Published: (2024)
by: Zhuang, Jingyu, et al.
Published: (2024)
Causal-Adapter: Taming Text-to-Image Diffusion for Faithful Counterfactual Generation
by: Tong, Lei, et al.
Published: (2025)
by: Tong, Lei, et al.
Published: (2025)
In-Context LoRA for Diffusion Transformers
by: Huang, Lianghua, et al.
Published: (2024)
by: Huang, Lianghua, et al.
Published: (2024)
ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling
by: Mao, Chaojie, et al.
Published: (2025)
by: Mao, Chaojie, et al.
Published: (2025)
Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models
by: Huang, Yupan, et al.
Published: (2023)
by: Huang, Yupan, et al.
Published: (2023)
PlanarSplatting: Accurate Planar Surface Reconstruction in 3 Minutes
by: Tan, Bin, et al.
Published: (2024)
by: Tan, Bin, et al.
Published: (2024)
Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis
by: Cao, Hengyuan, et al.
Published: (2025)
by: Cao, Hengyuan, et al.
Published: (2025)
DiffDoctor: Diagnosing Image Diffusion Models Before Treating
by: Wang, Yiyang, et al.
Published: (2025)
by: Wang, Yiyang, et al.
Published: (2025)
Reliable and Efficient Concept Erasure of Text-to-Image Diffusion Models
by: Gong, Chao, et al.
Published: (2024)
by: Gong, Chao, et al.
Published: (2024)
You Only Sample Once: Taming One-Step Text-to-Image Synthesis by Self-Cooperative Diffusion GANs
by: Luo, Yihong, et al.
Published: (2024)
by: Luo, Yihong, et al.
Published: (2024)
Learning Visual Generative Priors without Text
by: Ma, Shuailei, et al.
Published: (2024)
by: Ma, Shuailei, et al.
Published: (2024)
Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
by: Zhou, Chao, et al.
Published: (2025)
by: Zhou, Chao, et al.
Published: (2025)
Seek for Incantations: Towards Accurate Text-to-Image Diffusion Synthesis through Prompt Engineering
by: Yu, Chang, et al.
Published: (2024)
by: Yu, Chang, et al.
Published: (2024)
SpatialLock: Precise Spatial Control in Text-to-Image Synthesis
by: Liu, Biao, et al.
Published: (2025)
by: Liu, Biao, et al.
Published: (2025)
FlashFace: Human Image Personalization with High-fidelity Identity Preservation
by: Zhang, Shilong, et al.
Published: (2024)
by: Zhang, Shilong, et al.
Published: (2024)
Accurate Compression of Text-to-Image Diffusion Models via Vector Quantization
by: Egiazarian, Vage, et al.
Published: (2024)
by: Egiazarian, Vage, et al.
Published: (2024)
UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation
by: Guo, Qin, et al.
Published: (2025)
by: Guo, Qin, et al.
Published: (2025)
TACO: Taming Diffusion for in-the-wild Video Amodal Completion
by: Lu, Ruijie, et al.
Published: (2025)
by: Lu, Ruijie, et al.
Published: (2025)
Taming Diffusion Prior for Image Super-Resolution with Domain Shift SDEs
by: Cui, Qinpeng, et al.
Published: (2024)
by: Cui, Qinpeng, et al.
Published: (2024)
Calligrapher: Freestyle Text Image Customization
by: Ma, Yue, et al.
Published: (2025)
by: Ma, Yue, et al.
Published: (2025)
Taming Diffusion Models for Image Restoration: A Review
by: Luo, Ziwei, et al.
Published: (2024)
by: Luo, Ziwei, et al.
Published: (2024)
Rectified Diffusion Guidance for Conditional Generation
by: Xia, Mengfei, et al.
Published: (2024)
by: Xia, Mengfei, et al.
Published: (2024)
AI-T2I: Aggregating-and-Isolating Cross-Attention to Diffusion Models for Text-to-Image Synthesis
by: Cao, Shipeng, et al.
Published: (2026)
by: Cao, Shipeng, et al.
Published: (2026)
DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
by: Zhou, Dewei, et al.
Published: (2025)
by: Zhou, Dewei, et al.
Published: (2025)
GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction
by: Li, Jiahe, et al.
Published: (2025)
by: Li, Jiahe, et al.
Published: (2025)
From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models
by: Xiao, Changming, et al.
Published: (2023)
by: Xiao, Changming, et al.
Published: (2023)
InstructEngine: Instruction-driven Text-to-Image Alignment
by: Lu, Xingyu, et al.
Published: (2025)
by: Lu, Xingyu, et al.
Published: (2025)
LightningDrag: Lightning Fast and Accurate Drag-based Image Editing Emerging from Videos
by: Shi, Yujun, et al.
Published: (2024)
by: Shi, Yujun, et al.
Published: (2024)
Controllable Generation with Text-to-Image Diffusion Models: A Survey
by: Cao, Pu, et al.
Published: (2024)
by: Cao, Pu, et al.
Published: (2024)
Similar Items
-
Mimir: Improving Video Diffusion Models for Precise Text Understanding
by: Tan, Shuai, et al.
Published: (2024) -
Learning Disentangled Identifiers for Action-Customized Text-to-Image Generation
by: Huang, Siteng, et al.
Published: (2023) -
Check, Locate, Rectify: A Training-Free Layout Calibration System for Text-to-Image Generation
by: Gong, Biao, et al.
Published: (2023) -
ACE: All-round Creator and Editor Following Instructions via Diffusion Transformer
by: Han, Zhen, et al.
Published: (2024) -
Zero-shot Image Editing with Reference Imitation
by: Chen, Xi, et al.
Published: (2024)