D$^2$iT: Dynamic Diffusion Transformer for Accurate Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Jia, Weinan, Huang, Mengqi, Chen, Nan, Zhang, Lei, Mao, Zhendong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NativeTok: Native Visual Tokenization for Improved Image Generation
by: Wu, Bin, et al.
Published: (2026)
by: Wu, Bin, et al.
Published: (2026)
LongAnimation: Long Animation Generation with Dynamic Global-Local Memory
by: Chen, Nan, et al.
Published: (2025)
by: Chen, Nan, et al.
Published: (2025)
DSE-GAN: Dynamic Semantic Evolution Generative Adversarial Network for Text-to-Image Generation
by: Huang, Mengqi, et al.
Published: (2022)
by: Huang, Mengqi, et al.
Published: (2022)
HDGlyph: A Hierarchical Disentangled Glyph-Based Framework for Long-Tail Text Rendering in Diffusion Models
by: Zhuang, Shuhan, et al.
Published: (2025)
by: Zhuang, Shuhan, et al.
Published: (2025)
CustomContrast: A Multilevel Contrastive Perspective For Subject-Driven Text-to-Image Customization
by: Chen, Nan, et al.
Published: (2024)
by: Chen, Nan, et al.
Published: (2024)
MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
by: Jia, Weinan, et al.
Published: (2025)
by: Jia, Weinan, et al.
Published: (2025)
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
by: Wu, Bin, et al.
Published: (2026)
by: Wu, Bin, et al.
Published: (2026)
LayerEdit: Disentangled Multi-Object Editing via Conflict-Aware Multi-Layer Learning
by: Fu, Fengyi, et al.
Published: (2025)
by: Fu, Fengyi, et al.
Published: (2025)
RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image Customization
by: Huang, Mengqi, et al.
Published: (2024)
by: Huang, Mengqi, et al.
Published: (2024)
RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models
by: Lin, Yijing, et al.
Published: (2025)
by: Lin, Yijing, et al.
Published: (2025)
RealCustom++: Representing Images as Real Textual Word for Real-Time Customization
by: Mao, Zhendong, et al.
Published: (2024)
by: Mao, Zhendong, et al.
Published: (2024)
Gradual Residuals Alignment: A Dual-Stream Framework for GAN Inversion and Image Attribute Editing
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video Customization
by: Wang, Wenchuan, et al.
Published: (2025)
by: Wang, Wenchuan, et al.
Published: (2025)
Stream-T1: Test-Time Scaling for Streaming Video Generation
by: Tu, Yijing, et al.
Published: (2026)
by: Tu, Yijing, et al.
Published: (2026)
CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers
by: Chen, Weidong, et al.
Published: (2026)
by: Chen, Weidong, et al.
Published: (2026)
ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation
by: Zhao, Tianchen, et al.
Published: (2024)
by: Zhao, Tianchen, et al.
Published: (2024)
Dual-path Collaborative Generation Network for Emotional Video Captioning
by: Ye, Cheng, et al.
Published: (2024)
by: Ye, Cheng, et al.
Published: (2024)
DesignDiffusion: High-Quality Text-to-Design Image Generation with Diffusion Models
by: Wang, Zhendong, et al.
Published: (2025)
by: Wang, Zhendong, et al.
Published: (2025)
Diffusion Noise Feature: Accurate and Fast Generated Image Detection
by: Zhang, Yichi, et al.
Published: (2023)
by: Zhang, Yichi, et al.
Published: (2023)
Flash Window Attention: speedup the attention computation for Swin Transformer
by: Zhang, Zhendong
Published: (2025)
by: Zhang, Zhendong
Published: (2025)
2L3: Lifting Imperfect Generated 2D Images into Accurate 3D
by: Chen, Yizheng, et al.
Published: (2024)
by: Chen, Yizheng, et al.
Published: (2024)
Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
Med-Art: Diffusion Transformer for 2D Medical Text-to-Image Generation
by: Guo, Changlu, et al.
Published: (2025)
by: Guo, Changlu, et al.
Published: (2025)
Noise-Informed Diffusion-Generated Image Detection with Anomaly Attention
by: Guan, Weinan, et al.
Published: (2025)
by: Guan, Weinan, et al.
Published: (2025)
AccDiffusion: An Accurate Method for Higher-Resolution Image Generation
by: Lin, Zhihang, et al.
Published: (2024)
by: Lin, Zhihang, et al.
Published: (2024)
iVPT: Improving Task-relevant Information Sharing in Visual Prompt Tuning by Cross-layer Dynamic Connection
by: Zhou, Nan, et al.
Published: (2024)
by: Zhou, Nan, et al.
Published: (2024)
Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
Seek for Incantations: Towards Accurate Text-to-Image Diffusion Synthesis through Prompt Engineering
by: Yu, Chang, et al.
Published: (2024)
by: Yu, Chang, et al.
Published: (2024)
VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control
by: Wu, Shaojin, et al.
Published: (2024)
by: Wu, Shaojin, et al.
Published: (2024)
GeoDiffusion: A Training-Free Framework for Accurate 3D Geometric Conditioning in Image Generation
by: Mueller, Phillip, et al.
Published: (2025)
by: Mueller, Phillip, et al.
Published: (2025)
Acceleration Multiple Heads Decoding for LLM via Dynamic Tree Attention
by: Zhang, Zhendong
Published: (2025)
by: Zhang, Zhendong
Published: (2025)
EFDiT: Efficient Fine-grained Image Generation Using Diffusion Transformer Models
by: Wang, Kun, et al.
Published: (2025)
by: Wang, Kun, et al.
Published: (2025)
Dynamic Diffusion Transformer
by: Zhao, Wangbo, et al.
Published: (2024)
by: Zhao, Wangbo, et al.
Published: (2024)
Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer
by: Wu, Shuang, et al.
Published: (2024)
by: Wu, Shuang, et al.
Published: (2024)
Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers
by: Ma, Xin, et al.
Published: (2025)
by: Ma, Xin, et al.
Published: (2025)
Learning on Less: Constraining Pre-trained Model Learning for Generalizable Diffusion-Generated Image Detection
by: Chen, Yingjian, et al.
Published: (2024)
by: Chen, Yingjian, et al.
Published: (2024)
iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation
by: Fu, Zhoujie, et al.
Published: (2025)
by: Fu, Zhoujie, et al.
Published: (2025)
Dynamic 2D Gaussians: Geometrically Accurate Radiance Fields for Dynamic Objects
by: Zhang, Shuai, et al.
Published: (2024)
by: Zhang, Shuai, et al.
Published: (2024)
Latte: Latent Diffusion Transformer for Video Generation
by: Ma, Xin, et al.
Published: (2024)
by: Ma, Xin, et al.
Published: (2024)
GenTron: Diffusion Transformers for Image and Video Generation
by: Chen, Shoufa, et al.
Published: (2023)
by: Chen, Shoufa, et al.
Published: (2023)
Similar Items
-
NativeTok: Native Visual Tokenization for Improved Image Generation
by: Wu, Bin, et al.
Published: (2026) -
LongAnimation: Long Animation Generation with Dynamic Global-Local Memory
by: Chen, Nan, et al.
Published: (2025) -
DSE-GAN: Dynamic Semantic Evolution Generative Adversarial Network for Text-to-Image Generation
by: Huang, Mengqi, et al.
Published: (2022) -
HDGlyph: A Hierarchical Disentangled Glyph-Based Framework for Long-Tail Text Rendering in Diffusion Models
by: Zhuang, Shuhan, et al.
Published: (2025) -
CustomContrast: A Multilevel Contrastive Perspective For Subject-Driven Text-to-Image Customization
by: Chen, Nan, et al.
Published: (2024)