Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators
Fuente:
arXiv
Saved in:
| Main Authors: | Pu, Yifan, Xia, Zhuofan, Guo, Jiayi, Han, Dongchen, Li, Qixiu, Li, Duo, Yuan, Yuhui, Li, Ji, Han, Yizeng, Song, Shiji, Huang, Gao, Li, Xiu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging the Divide: Reconsidering Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2024)
by: Han, Dongchen, et al.
Published: (2024)
GSVA: Generalized Segmentation via Multimodal Large Language Models
by: Xia, Zhuofan, et al.
Published: (2023)
by: Xia, Zhuofan, et al.
Published: (2023)
Agent Attention: On the Integration of Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2023)
by: Han, Dongchen, et al.
Published: (2023)
Demystify Mamba in Vision: A Linear Attention Perspective
by: Han, Dongchen, et al.
Published: (2024)
by: Han, Dongchen, et al.
Published: (2024)
GRA: Detecting Oriented Objects through Group-wise Rotating and Attention
by: Wang, Jiangshan, et al.
Published: (2024)
by: Wang, Jiangshan, et al.
Published: (2024)
Latency-aware Unified Dynamic Networks for Efficient Image Recognition
by: Han, Yizeng, et al.
Published: (2023)
by: Han, Yizeng, et al.
Published: (2023)
Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise Differentials
by: Pu, Yifan, et al.
Published: (2025)
by: Pu, Yifan, et al.
Published: (2025)
Step by Step Network
by: Han, Dongchen, et al.
Published: (2025)
by: Han, Dongchen, et al.
Published: (2025)
Vision Transformers are Circulant Attention Learners
by: Han, Dongchen, et al.
Published: (2025)
by: Han, Dongchen, et al.
Published: (2025)
EfficientTrain++: Generalized Curriculum Learning for Efficient Visual Backbone Training
by: Wang, Yulin, et al.
Published: (2024)
by: Wang, Yulin, et al.
Published: (2024)
OStr-DARTS: Differentiable Neural Architecture Search based on Operation Strength
by: Yang, Le, et al.
Published: (2024)
by: Yang, Le, et al.
Published: (2024)
Emulating Human-like Adaptive Vision for Efficient and Flexible Machine Visual Perception
by: Wang, Yulin, et al.
Published: (2025)
by: Wang, Yulin, et al.
Published: (2025)
DyFADet: Dynamic Feature Aggregation for Temporal Action Detection
by: Yang, Le, et al.
Published: (2024)
by: Yang, Le, et al.
Published: (2024)
Few-Step Distillation for Text-to-Image Generation: A Practical Guide
by: Pu, Yifan, et al.
Published: (2025)
by: Pu, Yifan, et al.
Published: (2025)
DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution
by: Yue, Yang, et al.
Published: (2024)
by: Yue, Yang, et al.
Published: (2024)
Dynamic Diffusion Transformer
by: Zhao, Wangbo, et al.
Published: (2024)
by: Zhao, Wangbo, et al.
Published: (2024)
Linear-Time Global Visual Modeling without Explicit Attention
by: He, Ruize, et al.
Published: (2026)
by: He, Ruize, et al.
Published: (2026)
DyDiT++: Diffusion Transformers with Timestep and Spatial Dynamics for Efficient Visual Generation
by: Zhao, Wangbo, et al.
Published: (2025)
by: Zhao, Wangbo, et al.
Published: (2025)
Advancing Generalization in PINNs through Latent-Space Representations
by: Wang, Honghui, et al.
Published: (2024)
by: Wang, Honghui, et al.
Published: (2024)
Exploring contextual modeling with linear complexity for point cloud segmentation
by: Chng, Yong Xien, et al.
Published: (2024)
by: Chng, Yong Xien, et al.
Published: (2024)
LegalReasoner: Step-wised Verification-Correction for Legal Judgment Reasoning
by: Shi, Weijie, et al.
Published: (2025)
by: Shi, Weijie, et al.
Published: (2025)
RiboDiffusion: Tertiary Structure-based RNA Inverse Folding with Generative Diffusion Models
by: Huang, Han, et al.
Published: (2024)
by: Huang, Han, et al.
Published: (2024)
$C^1$-robust homoclinic tangencies
by: Li, Dongchen
Published: (2024)
by: Li, Dongchen
Published: (2024)
Blender-producing mechanisms and a dichotomy for local dynamics for heterodimensional cycles
by: Li, Dongchen
Published: (2024)
by: Li, Dongchen
Published: (2024)
SimPro: A Simple Probabilistic Framework Towards Realistic Long-Tailed Semi-Supervised Learning
by: Du, Chaoqun, et al.
Published: (2024)
by: Du, Chaoqun, et al.
Published: (2024)
Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise Attention
by: Li, Peng, et al.
Published: (2024)
by: Li, Peng, et al.
Published: (2024)
Revisiting Non-Autoregressive Transformers for Efficient Image Synthesis
by: Ni, Zanlin, et al.
Published: (2024)
by: Ni, Zanlin, et al.
Published: (2024)
DRM: Diffusion-based Reward Model With Step-wise Guidance
by: Zhang, Jaxon, et al.
Published: (2026)
by: Zhang, Jaxon, et al.
Published: (2026)
SPARK: Igniting Communication-Efficient Decentralized Learning via Stage-wise Projected NTK and Accelerated Regularization
by: Xia, Li
Published: (2025)
by: Xia, Li
Published: (2025)
UniTTA: Unified Benchmark and Versatile Framework Towards Realistic Test-Time Adaptation
by: Du, Chaoqun, et al.
Published: (2024)
by: Du, Chaoqun, et al.
Published: (2024)
Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
by: Li, Pengxiang, et al.
Published: (2025)
by: Li, Pengxiang, et al.
Published: (2025)
Elastic Diffusion Transformer
by: Wang, Jiangshan, et al.
Published: (2026)
by: Wang, Jiangshan, et al.
Published: (2026)
SITCOM: Step-wise Triple-Consistent Diffusion Sampling for Inverse Problems
by: Alkhouri, Ismail, et al.
Published: (2024)
by: Alkhouri, Ismail, et al.
Published: (2024)
ChWDTA: Channel-wise Wavelet-Domain Transformer Attention and Entropy Modeling for Learned Image Compression
by: Fu, Haisheng, et al.
Published: (2026)
by: Fu, Haisheng, et al.
Published: (2026)
Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization
by: Liang, Zhanhao, et al.
Published: (2024)
by: Liang, Zhanhao, et al.
Published: (2024)
Towards 3D-Aware Video Diffusion Models: Render-Free Human Motion Control with Mesh Tokenization
by: Liang, Jingyun, et al.
Published: (2026)
by: Liang, Jingyun, et al.
Published: (2026)
EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction
by: Cai, Han, et al.
Published: (2022)
by: Cai, Han, et al.
Published: (2022)
Meta-Semi: A Meta-learning Approach for Semi-supervised Learning
by: Wang, Yulin, et al.
Published: (2020)
by: Wang, Yulin, et al.
Published: (2020)
SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration
by: Wen, Zhuofan, et al.
Published: (2026)
by: Wen, Zhuofan, et al.
Published: (2026)
GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer
by: Jia, Ding, et al.
Published: (2024)
by: Jia, Ding, et al.
Published: (2024)
Similar Items
-
Bridging the Divide: Reconsidering Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2024) -
GSVA: Generalized Segmentation via Multimodal Large Language Models
by: Xia, Zhuofan, et al.
Published: (2023) -
Agent Attention: On the Integration of Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2023) -
Demystify Mamba in Vision: A Linear Attention Perspective
by: Han, Dongchen, et al.
Published: (2024) -
GRA: Detecting Oriented Objects through Group-wise Rotating and Attention
by: Wang, Jiangshan, et al.
Published: (2024)