Rethinking Global Text Conditioning in Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Starodubcev, Nikita, Pakhomov, Daniil, Wu, Zongze, Drobyshevskiy, Ilya, Liu, Yuchen, Wang, Zhonghao, Zhou, Yuqian, Lin, Zhe, Baranchuk, Dmitry |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Registers Matter for Pixel-Space Diffusion Transformers
by: Starodubcev, Nikita, et al.
Published: (2026)
by: Starodubcev, Nikita, et al.
Published: (2026)
Scale-wise Distillation of Diffusion Models
by: Starodubcev, Nikita, et al.
Published: (2025)
by: Starodubcev, Nikita, et al.
Published: (2025)
Your Student is Better Than Expected: Adaptive Teacher-Student Collaboration for Text-Conditional Diffusion Models
by: Starodubcev, Nikita, et al.
Published: (2023)
by: Starodubcev, Nikita, et al.
Published: (2023)
Invertible Consistency Distillation for Text-Guided Image Editing in Around 7 Steps
by: Starodubcev, Nikita, et al.
Published: (2024)
by: Starodubcev, Nikita, et al.
Published: (2024)
TurboFill: Adapting Few-step Text-to-image Model for Fast Image Inpainting
by: Xie, Liangbin, et al.
Published: (2025)
by: Xie, Liangbin, et al.
Published: (2025)
Accurate Compression of Text-to-Image Diffusion Models via Vector Quantization
by: Egiazarian, Vage, et al.
Published: (2024)
by: Egiazarian, Vage, et al.
Published: (2024)
Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis
by: Voronov, Anton, et al.
Published: (2024)
by: Voronov, Anton, et al.
Published: (2024)
Inverse Bridge Matching Distillation
by: Gushchin, Nikita, et al.
Published: (2025)
by: Gushchin, Nikita, et al.
Published: (2025)
CasTex: Cascaded Text-to-Texture Synthesis via Explicit Texture Maps and Physically-Based Shading
by: Aliev, Mishan, et al.
Published: (2025)
by: Aliev, Mishan, et al.
Published: (2025)
MADrive: Memory-Augmented Driving Scene Modeling
by: Karpikova, Polina, et al.
Published: (2025)
by: Karpikova, Polina, et al.
Published: (2025)
Alchemist: Turning Public Text-to-Image Data into Generative Gold
by: Startsev, Valerii, et al.
Published: (2025)
by: Startsev, Valerii, et al.
Published: (2025)
Revisiting Autoregressive Models for Generative Image Classification
by: Sudakov, Ilia, et al.
Published: (2026)
by: Sudakov, Ilia, et al.
Published: (2026)
ZipIR: Latent Pyramid Diffusion Transformer for High-Resolution Image Restoration
by: Yu, Yongsheng, et al.
Published: (2025)
by: Yu, Yongsheng, et al.
Published: (2025)
Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
by: Zhang, Shilong, et al.
Published: (2025)
by: Zhang, Shilong, et al.
Published: (2025)
Gradient as Conditions: Rethinking HOG for All-in-one Image Restoration
by: Wu, Jiawei, et al.
Published: (2025)
by: Wu, Jiawei, et al.
Published: (2025)
PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement
by: Zheng, Haitian, et al.
Published: (2025)
by: Zheng, Haitian, et al.
Published: (2025)
Enhancing Text-to-Image Diffusion Transformer via Split-Text Conditioning
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
by: Ju, Xuan, et al.
Published: (2025)
by: Ju, Xuan, et al.
Published: (2025)
Training-Free Anomaly Generation via Dual-Attention Enhancement in Diffusion Model
by: Zuo, Zuo, et al.
Published: (2025)
by: Zuo, Zuo, et al.
Published: (2025)
Generative Image Layer Decomposition with Visual Effects
by: Yang, Jinrui, et al.
Published: (2024)
by: Yang, Jinrui, et al.
Published: (2024)
Object-level Scene Deocclusion
by: Liu, Zhengzhe, et al.
Published: (2024)
by: Liu, Zhengzhe, et al.
Published: (2024)
Rethinking Query-based Transformer for Continual Image Segmentation
by: Zhu, Yuchen, et al.
Published: (2025)
by: Zhu, Yuchen, et al.
Published: (2025)
Lazy Diffusion Transformer for Interactive Image Editing
by: Nitzan, Yotam, et al.
Published: (2024)
by: Nitzan, Yotam, et al.
Published: (2024)
Diffusion in Diffusion: Cyclic One-Way Diffusion for Text-Vision-Conditioned Generation
by: Wang, Ruoyu, et al.
Published: (2023)
by: Wang, Ruoyu, et al.
Published: (2023)
Rethinking Cross-Layer Information Routing in Diffusion Transformers
by: Xu, Chao, et al.
Published: (2026)
by: Xu, Chao, et al.
Published: (2026)
Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers
by: Lv, Zhengyao, et al.
Published: (2025)
by: Lv, Zhengyao, et al.
Published: (2025)
Local Conditional Controlling for Text-to-Image Diffusion Models
by: Zhao, Yibo, et al.
Published: (2023)
by: Zhao, Yibo, et al.
Published: (2023)
ForgeDreamer: Industrial Text-to-3D Generation with Multi-Expert LoRA and Cross-View Hypergraph
by: Cai, Junhao, et al.
Published: (2026)
by: Cai, Junhao, et al.
Published: (2026)
SCoRD: Subject-Conditional Relation Detection with Text-Augmented Data
by: Yang, Ziyan, et al.
Published: (2023)
by: Yang, Ziyan, et al.
Published: (2023)
Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability,Reproducibility, and Practicality
by: Zhang, Tianle, et al.
Published: (2024)
by: Zhang, Tianle, et al.
Published: (2024)
ACDiT: Interpolating Autoregressive Conditional Modeling and Diffusion Transformer
by: Hu, Jinyi, et al.
Published: (2024)
by: Hu, Jinyi, et al.
Published: (2024)
FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers
by: Zhang, Yanbing, et al.
Published: (2025)
by: Zhang, Yanbing, et al.
Published: (2025)
UltraImage: Rethinking Resolution Extrapolation in Image Diffusion Transformers
by: Zhao, Min, et al.
Published: (2025)
by: Zhao, Min, et al.
Published: (2025)
A3D: Does Diffusion Dream about 3D Alignment?
by: Ignatyev, Savva, et al.
Published: (2024)
by: Ignatyev, Savva, et al.
Published: (2024)
Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion
by: Zhou, Zhenghong, et al.
Published: (2026)
by: Zhou, Zhenghong, et al.
Published: (2026)
Using predefined vector systems to speed up neural network multimillion class classification
by: Gabdullin, Nikita, et al.
Published: (2026)
by: Gabdullin, Nikita, et al.
Published: (2026)
Mixture of Efficient Diffusion Experts Through Automatic Interval and Sub-Network Selection
by: Ganjdanesh, Alireza, et al.
Published: (2024)
by: Ganjdanesh, Alireza, et al.
Published: (2024)
Rethinking Garment Conditioning in Diffusion-based Virtual Try-On
by: Na, Kihyun, et al.
Published: (2025)
by: Na, Kihyun, et al.
Published: (2025)
EasyText: Controllable Diffusion Transformer for Multilingual Text Rendering
by: Lu, Runnan, et al.
Published: (2025)
by: Lu, Runnan, et al.
Published: (2025)
Object-Conditioned Energy-Based Attention Map Alignment in Text-to-Image Diffusion Models
by: Zhang, Yasi, et al.
Published: (2024)
by: Zhang, Yasi, et al.
Published: (2024)
Similar Items
-
Registers Matter for Pixel-Space Diffusion Transformers
by: Starodubcev, Nikita, et al.
Published: (2026) -
Scale-wise Distillation of Diffusion Models
by: Starodubcev, Nikita, et al.
Published: (2025) -
Your Student is Better Than Expected: Adaptive Teacher-Student Collaboration for Text-Conditional Diffusion Models
by: Starodubcev, Nikita, et al.
Published: (2023) -
Invertible Consistency Distillation for Text-Guided Image Editing in Around 7 Steps
by: Starodubcev, Nikita, et al.
Published: (2024) -
TurboFill: Adapting Few-step Text-to-image Model for Fast Image Inpainting
by: Xie, Liangbin, et al.
Published: (2025)