Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Feng, Zhang, Shiwei, Wang, Xiaofeng, Wei, Yujie, Qiu, Haonan, Zhao, Yuzhong, Zhang, Yingya, Ye, Qixiang, Wan, Fang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale Fusion
by: Qiu, Haonan, et al.
Published: (2024)
by: Qiu, Haonan, et al.
Published: (2024)
DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control
by: Wei, Yujie, et al.
Published: (2024)
by: Wei, Yujie, et al.
Published: (2024)
PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic Degradation
by: Li, Hengjia, et al.
Published: (2024)
by: Li, Hengjia, et al.
Published: (2024)
Thinking with Images via Self-Calling Agent
by: Yang, Wenxi, et al.
Published: (2025)
by: Yang, Wenxi, et al.
Published: (2025)
DreamRelation: Relation-Centric Video Customization
by: Wei, Yujie, et al.
Published: (2025)
by: Wei, Yujie, et al.
Published: (2025)
Evaluation of Text-to-Video Generation Models: A Dynamics Perspective
by: Liao, Mingxiang, et al.
Published: (2024)
by: Liao, Mingxiang, et al.
Published: (2024)
Timestep-Aware Correction for Quantized Diffusion Models
by: Yao, Yuzhe, et al.
Published: (2024)
by: Yao, Yuzhe, et al.
Published: (2024)
TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation
by: Chen, Zhekai, et al.
Published: (2025)
by: Chen, Zhekai, et al.
Published: (2025)
DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning
by: Wei, Yujie, et al.
Published: (2026)
by: Wei, Yujie, et al.
Published: (2026)
Replace Anyone in Videos
by: Wang, Xiang, et al.
Published: (2024)
by: Wang, Xiang, et al.
Published: (2024)
DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution
by: Zhao, Yuzhong, et al.
Published: (2024)
by: Zhao, Yuzhong, et al.
Published: (2024)
UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation
by: Wang, Xiang, et al.
Published: (2024)
by: Wang, Xiang, et al.
Published: (2024)
Accelerating Diffusion-based Video Editing via Heterogeneous Caching: Beyond Full Computing at Sampled Denoising Timestep
by: Liu, Tianyi, et al.
Published: (2026)
by: Liu, Tianyi, et al.
Published: (2026)
DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models
by: Ma, Yifeng, et al.
Published: (2023)
by: Ma, Yifeng, et al.
Published: (2023)
CC-Diff: Enhancing Contextual Coherence in Remote Sensing Image Synthesis
by: Zhang, Mu, et al.
Published: (2024)
by: Zhang, Mu, et al.
Published: (2024)
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
by: Wang, Xiang, et al.
Published: (2025)
by: Wang, Xiang, et al.
Published: (2025)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
by: Cai, Lingling, et al.
Published: (2024)
by: Cai, Lingling, et al.
Published: (2024)
TASR: Timestep-Aware Diffusion Model for Image Super-Resolution
by: Lin, Qinwei, et al.
Published: (2024)
by: Lin, Qinwei, et al.
Published: (2024)
ControlCap: Controllable Region-level Captioning
by: Zhao, Yuzhong, et al.
Published: (2024)
by: Zhao, Yuzhong, et al.
Published: (2024)
Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning
by: Lin, Wang, et al.
Published: (2025)
by: Lin, Wang, et al.
Published: (2025)
Compute Only 16 Tokens in One Timestep: Accelerating Diffusion Transformers with Cluster-Driven Feature Caching
by: Zheng, Zhixin, et al.
Published: (2025)
by: Zheng, Zhixin, et al.
Published: (2025)
Taming Consistency Distillation for Accelerated Human Image Animation
by: Wang, Xiang, et al.
Published: (2025)
by: Wang, Xiang, et al.
Published: (2025)
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
by: Wang, Xiaofeng, et al.
Published: (2024)
by: Wang, Xiaofeng, et al.
Published: (2024)
Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
by: Wei, Yujie, et al.
Published: (2025)
by: Wei, Yujie, et al.
Published: (2025)
Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
by: Pan, Kaihang, et al.
Published: (2025)
by: Pan, Kaihang, et al.
Published: (2025)
Redefining Temporal Modeling in Video Diffusion: The Vectorized Timestep Approach
by: Liu, Yaofang, et al.
Published: (2024)
by: Liu, Yaofang, et al.
Published: (2024)
EvolveDirector: Approaching Advanced Text-to-Image Generation with Large Vision-Language Models
by: Zhao, Rui, et al.
Published: (2024)
by: Zhao, Rui, et al.
Published: (2024)
Timestep-Aware Block Masking for Efficient Diffusion Model Inference
by: He, Haodong, et al.
Published: (2026)
by: He, Haodong, et al.
Published: (2026)
S&D Messenger: Exchanging Semantic and Domain Knowledge for Generic Semi-Supervised Medical Image Segmentation
by: Zhang, Qixiang, et al.
Published: (2024)
by: Zhang, Qixiang, et al.
Published: (2024)
Tuning Timestep-Distilled Diffusion Model Using Pairwise Sample Optimization
by: Miao, Zichen, et al.
Published: (2024)
by: Miao, Zichen, et al.
Published: (2024)
Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation
by: Liu, Yaofang, et al.
Published: (2025)
by: Liu, Yaofang, et al.
Published: (2025)
MagCache: Fast Video Generation with Magnitude-Aware Cache
by: Ma, Zehong, et al.
Published: (2025)
by: Ma, Zehong, et al.
Published: (2025)
Characterizing Motion Encoding in Video Diffusion Timesteps
by: Baherwani, Vatsal, et al.
Published: (2025)
by: Baherwani, Vatsal, et al.
Published: (2025)
CLIP-guided Prototype Modulating for Few-shot Action Recognition
by: Wang, Xiang, et al.
Published: (2023)
by: Wang, Xiang, et al.
Published: (2023)
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
by: Qiu, Jihao, et al.
Published: (2026)
by: Qiu, Jihao, et al.
Published: (2026)
DiCache: Let Diffusion Model Determine Its Own Cache
by: Bu, Jiazi, et al.
Published: (2025)
by: Bu, Jiazi, et al.
Published: (2025)
Towards More Accurate Diffusion Model Acceleration with A Timestep Tuner
by: Xia, Mengfei, et al.
Published: (2023)
by: Xia, Mengfei, et al.
Published: (2023)
HybridStitch: Pixel and Timestep Level Model Stitching for Diffusion Acceleration
by: Sun, Desen, et al.
Published: (2026)
by: Sun, Desen, et al.
Published: (2026)
CTCal: Rethinking Text-to-Image Diffusion Models via Cross-Timestep Self-Calibration
by: Guo, Xiefan, et al.
Published: (2026)
by: Guo, Xiefan, et al.
Published: (2026)
Neurons: Emulating the Human Visual Cortex Improves Fidelity and Interpretability in fMRI-to-Video Reconstruction
by: Wang, Haonan, et al.
Published: (2025)
by: Wang, Haonan, et al.
Published: (2025)
Similar Items
-
FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale Fusion
by: Qiu, Haonan, et al.
Published: (2024) -
DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control
by: Wei, Yujie, et al.
Published: (2024) -
PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic Degradation
by: Li, Hengjia, et al.
Published: (2024) -
Thinking with Images via Self-Calling Agent
by: Yang, Wenxi, et al.
Published: (2025) -
DreamRelation: Relation-Centric Video Customization
by: Wei, Yujie, et al.
Published: (2025)