U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Tian, Yuchuan, Tu, Zhijun, Chen, Hanting, Hu, Jie, Xu, Chao, Wang, Yunhe |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
U-REPA: Aligning Diffusion U-Nets to ViTs
by: Tian, Yuchuan, et al.
Published: (2025)
by: Tian, Yuchuan, et al.
Published: (2025)
Unveiling Redundancy in Diffusion Transformers (DiTs): A Systematic Study
by: Sun, Xibo, et al.
Published: (2024)
by: Sun, Xibo, et al.
Published: (2024)
Instruct-IPT: All-in-One Image Processing Transformer via Weight Modulation
by: Tian, Yuchuan, et al.
Published: (2024)
by: Tian, Yuchuan, et al.
Published: (2024)
DiC: Rethinking Conv3x3 Designs in Diffusion Models
by: Tian, Yuchuan, et al.
Published: (2024)
by: Tian, Yuchuan, et al.
Published: (2024)
IPT-V2: Efficient Image Processing Transformer using Hierarchical Attentions
by: Tu, Zhijun, et al.
Published: (2024)
by: Tu, Zhijun, et al.
Published: (2024)
AnchorDiff: Training-Free Concept Grounding for MM-DiTs via Anchor-Based Graph Propagation
by: Zhang, Jian, et al.
Published: (2026)
by: Zhang, Jian, et al.
Published: (2026)
DiT4Edit: Diffusion Transformer for Image Editing
by: Feng, Kunyu, et al.
Published: (2024)
by: Feng, Kunyu, et al.
Published: (2024)
LaVin-DiT: Large Vision Diffusion Transformer
by: Wang, Zhaoqing, et al.
Published: (2024)
by: Wang, Zhaoqing, et al.
Published: (2024)
$Δ$-DiT: A Training-Free Acceleration Method Tailored for Diffusion Transformers
by: Chen, Pengtao, et al.
Published: (2024)
by: Chen, Pengtao, et al.
Published: (2024)
Untwisting RoPE: Frequency Control for Shared Attention in DiTs
by: Mikaeili, Aryan, et al.
Published: (2026)
by: Mikaeili, Aryan, et al.
Published: (2026)
MeshCraft: Exploring Efficient and Controllable Mesh Generation with Flow-based DiTs
by: He, Xianglong, et al.
Published: (2025)
by: He, Xianglong, et al.
Published: (2025)
S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation
by: Zhao, Lin, et al.
Published: (2026)
by: Zhao, Lin, et al.
Published: (2026)
PTQ4DiT: Post-training Quantization for Diffusion Transformers
by: Wu, Junyi, et al.
Published: (2024)
by: Wu, Junyi, et al.
Published: (2024)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
Rethinking Multi-Condition DiTs: Eliminating Redundant Attention via Position-Alignment and Keyword-Scoping
by: Zhou, Chao, et al.
Published: (2026)
by: Zhou, Chao, et al.
Published: (2026)
Inf-DiT: Upsampling Any-Resolution Image with Memory-Efficient Diffusion Transformer
by: Yang, Zhuoyi, et al.
Published: (2024)
by: Yang, Zhuoyi, et al.
Published: (2024)
Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
by: Li, Zhimin, et al.
Published: (2024)
by: Li, Zhimin, et al.
Published: (2024)
FP4DiT: Towards Effective Floating Point Quantization for Diffusion Transformers
by: Chen, Ruichen, et al.
Published: (2025)
by: Chen, Ruichen, et al.
Published: (2025)
Remix-DiT: Mixing Diffusion Transformers for Multi-Expert Denoising
by: Fang, Gongfan, et al.
Published: (2024)
by: Fang, Gongfan, et al.
Published: (2024)
Learning Quantized Adaptive Conditions for Diffusion Models
by: Liang, Yuchen, et al.
Published: (2024)
by: Liang, Yuchen, et al.
Published: (2024)
DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer
by: Lyu, Hengye, et al.
Published: (2026)
by: Lyu, Hengye, et al.
Published: (2026)
SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer
by: Zhu, Rui, et al.
Published: (2024)
by: Zhu, Rui, et al.
Published: (2024)
RobuQ: Pushing DiTs to W1.58A2 via Robust Activation Quantization
by: Yang, Kaicheng, et al.
Published: (2025)
by: Yang, Kaicheng, et al.
Published: (2025)
DiVE: DiT-based Video Generation with Enhanced Control
by: Jiang, Junpeng, et al.
Published: (2024)
by: Jiang, Junpeng, et al.
Published: (2024)
Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
by: Sun, Yasheng, et al.
Published: (2025)
by: Sun, Yasheng, et al.
Published: (2025)
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
by: Wang, Xiang, et al.
Published: (2025)
by: Wang, Xiang, et al.
Published: (2025)
DiT4SR: Taming Diffusion Transformer for Real-World Image Super-Resolution
by: Duan, Zheng-Peng, et al.
Published: (2025)
by: Duan, Zheng-Peng, et al.
Published: (2025)
VITON-DiT: Learning In-the-Wild Video Try-On from Human Dance Videos via Diffusion Transformers
by: Zheng, Jun, et al.
Published: (2024)
by: Zheng, Jun, et al.
Published: (2024)
LRQ-DiT: Log-Rotation Post-Training Quantization of Diffusion Transformers for Image and Video Generation
by: Yang, Lianwei, et al.
Published: (2025)
by: Yang, Lianwei, et al.
Published: (2025)
DiT-Air: Revisiting the Efficiency of Diffusion Model Architecture Design in Text to Image Generation
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
DiT-IC: Aligned Diffusion Transformer for Efficient Image Compression
by: Shi, Junqi, et al.
Published: (2026)
by: Shi, Junqi, et al.
Published: (2026)
What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing
by: Lin, Hangyu, et al.
Published: (2026)
by: Lin, Hangyu, et al.
Published: (2026)
HQ-DiT: Efficient Diffusion Transformer with FP4 Hybrid Quantization
by: Liu, Wenxuan, et al.
Published: (2024)
by: Liu, Wenxuan, et al.
Published: (2024)
LIPT: Latency-aware Image Processing Transformer
by: Qiao, Junbo, et al.
Published: (2024)
by: Qiao, Junbo, et al.
Published: (2024)
HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads
by: Xu, Yu, et al.
Published: (2024)
by: Xu, Yu, et al.
Published: (2024)
DiT-JSCC: Rethinking Deep JSCC with Diffusion Transformers and Semantic Representations
by: Tan, Kailin, et al.
Published: (2026)
by: Tan, Kailin, et al.
Published: (2026)
Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT Acceleration
by: Fang, Haipeng, et al.
Published: (2025)
by: Fang, Haipeng, et al.
Published: (2025)
ADP-DiT: Text-Guided Diffusion Transformer for Brain Image Generation in Alzheimer's Disease Progression
by: Lee, Juneyong, et al.
Published: (2026)
by: Lee, Juneyong, et al.
Published: (2026)
Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection
by: Li, Shufan, et al.
Published: (2025)
by: Li, Shufan, et al.
Published: (2025)
Similar Items
-
U-REPA: Aligning Diffusion U-Nets to ViTs
by: Tian, Yuchuan, et al.
Published: (2025) -
Unveiling Redundancy in Diffusion Transformers (DiTs): A Systematic Study
by: Sun, Xibo, et al.
Published: (2024) -
Instruct-IPT: All-in-One Image Processing Transformer via Weight Modulation
by: Tian, Yuchuan, et al.
Published: (2024) -
DiC: Rethinking Conv3x3 Designs in Diffusion Models
by: Tian, Yuchuan, et al.
Published: (2024) -
IPT-V2: Efficient Image Processing Transformer using Hierarchical Attentions
by: Tu, Zhijun, et al.
Published: (2024)