Not All Tokens Need 40 Steps: Heterogeneous Step Allocation in Diffusion Transformers for Efficient Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Chu, Ernie, Patel, Vishal M. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Face-to-Face: A Video Dataset for Multi-Person Interaction Modeling
by: Chu, Ernie, et al.
Published: (2026)
by: Chu, Ernie, et al.
Published: (2026)
StepAL: Step-aware Active Learning for Cataract Surgical Videos
by: Shah, Nisarg A., et al.
Published: (2025)
by: Shah, Nisarg A., et al.
Published: (2025)
SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
by: Vallaeys, Théophane, et al.
Published: (2025)
by: Vallaeys, Théophane, et al.
Published: (2025)
Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators
by: Pu, Yifan, et al.
Published: (2024)
by: Pu, Yifan, et al.
Published: (2024)
Clockwork Diffusion: Efficient Generation With Model-Step Distillation
by: Habibian, Amirhossein, et al.
Published: (2023)
by: Habibian, Amirhossein, et al.
Published: (2023)
Efficient One-Step Diffusion Restoration Model with Compact Token Compression and Linear Attention
by: Qiao, Bingtian, et al.
Published: (2026)
by: Qiao, Bingtian, et al.
Published: (2026)
Diffscaler: Enhancing the Generative Prowess of Diffusion Transformers
by: Nair, Nithin Gopalakrishnan, et al.
Published: (2024)
by: Nair, Nithin Gopalakrishnan, et al.
Published: (2024)
Step Saver: Predicting Minimum Denoising Steps for Diffusion Model Image Generation
by: Yu, Jean, et al.
Published: (2024)
by: Yu, Jean, et al.
Published: (2024)
Real-time One-Step Diffusion-based Expressive Portrait Videos Generation
by: Guo, Hanzhong, et al.
Published: (2024)
by: Guo, Hanzhong, et al.
Published: (2024)
Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization
by: He, Xiaoxuan, et al.
Published: (2026)
by: He, Xiaoxuan, et al.
Published: (2026)
VividFace: High-Quality and Efficient One-Step Diffusion For Video Face Enhancement
by: Zhang, Shulian, et al.
Published: (2025)
by: Zhang, Shulian, et al.
Published: (2025)
DOVE: Efficient One-Step Diffusion Model for Real-World Video Super-Resolution
by: Chen, Zheng, et al.
Published: (2025)
by: Chen, Zheng, et al.
Published: (2025)
Aligning Step-by-Step Instructional Diagrams to Video Demonstrations
by: Zhang, Jiahao, et al.
Published: (2023)
by: Zhang, Jiahao, et al.
Published: (2023)
VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step
by: Wang, Hanyang, et al.
Published: (2025)
by: Wang, Hanyang, et al.
Published: (2025)
Frame by Familiar Frame: Understanding Replication in Video Diffusion Models
by: Rahman, Aimon, et al.
Published: (2024)
by: Rahman, Aimon, et al.
Published: (2024)
Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders
by: Kumar, Amandeep, et al.
Published: (2026)
by: Kumar, Amandeep, et al.
Published: (2026)
Dreamguider: Improved Training free Diffusion-based Conditional Generation
by: Nair, Nithin Gopalakrishnan, et al.
Published: (2024)
by: Nair, Nithin Gopalakrishnan, et al.
Published: (2024)
Low-rank Adaptation-based All-Weather Removal for Autonomous Navigation
by: Rajagopalan, Sudarshan, et al.
Published: (2024)
by: Rajagopalan, Sudarshan, et al.
Published: (2024)
Diffusion Models Are Innate One-Step Generators
by: Zheng, Bowen, et al.
Published: (2024)
by: Zheng, Bowen, et al.
Published: (2024)
Phased One-Step Adversarial Equilibrium for Video Diffusion Models
by: Cheng, Jiaxiang, et al.
Published: (2025)
by: Cheng, Jiaxiang, et al.
Published: (2025)
AWRaCLe: All-Weather Image Restoration using Visual In-Context Learning
by: Rajagopalan, Sudarshan, et al.
Published: (2024)
by: Rajagopalan, Sudarshan, et al.
Published: (2024)
Need for Speed: Zero-Shot Depth Completion with Single-Step Diffusion
by: Gregorek, Jakub, et al.
Published: (2026)
by: Gregorek, Jakub, et al.
Published: (2026)
Think Before You Diffuse: Infusing Physical Rules into Video Diffusion
by: Zhang, Ke, et al.
Published: (2025)
by: Zhang, Ke, et al.
Published: (2025)
Step Differences in Instructional Video
by: Nagarajan, Tushar, et al.
Published: (2024)
by: Nagarajan, Tushar, et al.
Published: (2024)
EasyTune: Efficient Step-Aware Fine-Tuning for Diffusion-Based Motion Generation
by: Tan, Xiaofeng, et al.
Published: (2026)
by: Tan, Xiaofeng, et al.
Published: (2026)
UltraVSR: Achieving Ultra-Realistic Video Super-Resolution with Efficient One-Step Diffusion Space
by: Liu, Yong, et al.
Published: (2025)
by: Liu, Yong, et al.
Published: (2025)
Beyond Few-Step Inference: Accelerating Video Diffusion Transformer Model Serving with Inter-Request Caching Reuse
by: Liu, Hao, et al.
Published: (2026)
by: Liu, Hao, et al.
Published: (2026)
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
by: NextStep Team, et al.
Published: (2025)
by: NextStep Team, et al.
Published: (2025)
GenDeg: Diffusion-based Degradation Synthesis for Generalizable All-In-One Image Restoration
by: Rajagopalan, Sudarshan, et al.
Published: (2024)
by: Rajagopalan, Sudarshan, et al.
Published: (2024)
Two Steps Are All You Need: Efficient 3D Point Cloud Anomaly Detection with Consistency Models
by: A, Pranav, et al.
Published: (2026)
by: A, Pranav, et al.
Published: (2026)
Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
by: Zhao, Min, et al.
Published: (2026)
by: Zhao, Min, et al.
Published: (2026)
RestoreVAR: Visual Autoregressive Generation for All-in-One Image Restoration
by: Rajagopalan, Sudarshan, et al.
Published: (2025)
by: Rajagopalan, Sudarshan, et al.
Published: (2025)
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
Step by Step Network
by: Han, Dongchen, et al.
Published: (2025)
by: Han, Dongchen, et al.
Published: (2025)
Efficient Coarse-to-Fine Diffusion Models with Time Step Sequence Redistribution
by: Tai, Yu-Shan, et al.
Published: (2026)
by: Tai, Yu-Shan, et al.
Published: (2026)
One-Step Diffusion Transformer for Controllable Real-World Image Super-Resolution
by: Fang, Yushun, et al.
Published: (2025)
by: Fang, Yushun, et al.
Published: (2025)
Diffusion Adversarial Post-Training for One-Step Video Generation
by: Lin, Shanchuan, et al.
Published: (2025)
by: Lin, Shanchuan, et al.
Published: (2025)
Denoising Diffusion Step-aware Models
by: Yang, Shuai, et al.
Published: (2023)
by: Yang, Shuai, et al.
Published: (2023)
EFlow: Fast Few-Step Video Generator Training from Scratch via Efficient Solution Flow
by: Park, Dogyun, et al.
Published: (2026)
by: Park, Dogyun, et al.
Published: (2026)
You Only Need One Step: Fast Super-Resolution with Stable Diffusion via Scale Distillation
by: Noroozi, Mehdi, et al.
Published: (2024)
by: Noroozi, Mehdi, et al.
Published: (2024)
Similar Items
-
Face-to-Face: A Video Dataset for Multi-Person Interaction Modeling
by: Chu, Ernie, et al.
Published: (2026) -
StepAL: Step-aware Active Learning for Cataract Surgical Videos
by: Shah, Nisarg A., et al.
Published: (2025) -
SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
by: Vallaeys, Théophane, et al.
Published: (2025) -
Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators
by: Pu, Yifan, et al.
Published: (2024) -
Clockwork Diffusion: Efficient Generation With Model-Step Distillation
by: Habibian, Amirhossein, et al.
Published: (2023)