Temporal Aware Pruning for Efficient Diffusion-based Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Sheng, Sui, Yang, Ran, Junhao, Yuan, Bo, Dai, Yue, Tang, Xulong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TokenDance: Token-to-Token Music-to-Dance Generation with Bidirectional Mamba
by: Yang, Ziyue, et al.
Published: (2026)
by: Yang, Ziyue, et al.
Published: (2026)
EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent
by: Li, Jiaao, et al.
Published: (2025)
by: Li, Jiaao, et al.
Published: (2025)
Manifold-Aware Exploration for Reinforcement Learning in Video Generation
by: Zheng, Mingzhe, et al.
Published: (2026)
by: Zheng, Mingzhe, et al.
Published: (2026)
HierarchicalPrune: Position-Aware Compression for Large-Scale Diffusion Models
by: Kwon, Young D., et al.
Published: (2025)
by: Kwon, Young D., et al.
Published: (2025)
Cluster-based Video Summarization with Temporal Context Awareness
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
by: Qin, Jialong, et al.
Published: (2025)
by: Qin, Jialong, et al.
Published: (2025)
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
by: Tan, Zhentao, et al.
Published: (2024)
by: Tan, Zhentao, et al.
Published: (2024)
SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
by: Chen, Junsong, et al.
Published: (2025)
by: Chen, Junsong, et al.
Published: (2025)
Focused Forcing: Content-Aware Per-Frame KV Selection for Efficient Autoregressive Video Diffusion
by: Cai, Peiliang, et al.
Published: (2026)
by: Cai, Peiliang, et al.
Published: (2026)
TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model
by: Yang, Cheng, et al.
Published: (2025)
by: Yang, Cheng, et al.
Published: (2025)
Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning
by: Yue, Feng, et al.
Published: (2025)
by: Yue, Feng, et al.
Published: (2025)
Spatio-Temporal Token Pruning for Efficient High-Resolution GUI Agents
by: Xu, Zhou, et al.
Published: (2026)
by: Xu, Zhou, et al.
Published: (2026)
ELRT: Efficient Low-Rank Training for Compact Convolutional Neural Networks
by: Sui, Yang, et al.
Published: (2024)
by: Sui, Yang, et al.
Published: (2024)
ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models
by: Zhang, Pu, et al.
Published: (2025)
by: Zhang, Pu, et al.
Published: (2025)
Rethinking Video Tokenization: A Conditioned Diffusion-based Approach
by: Yang, Nianzu, et al.
Published: (2025)
by: Yang, Nianzu, et al.
Published: (2025)
Motion-Aware Caching for Efficient Autoregressive Video Generation
by: Xu, Jing, et al.
Published: (2026)
by: Xu, Jing, et al.
Published: (2026)
From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration
by: Shi, Jiaqi, et al.
Published: (2026)
by: Shi, Jiaqi, et al.
Published: (2026)
SpecSwin3D: Generating Hyperspectral Imagery from Multispectral Data via Transformer Networks
by: Sui, Tang, et al.
Published: (2025)
by: Sui, Tang, et al.
Published: (2025)
FADE: Frequency-Aware Diffusion Model Factorization for Video Editing
by: Zhu, Yixuan, et al.
Published: (2025)
by: Zhu, Yixuan, et al.
Published: (2025)
SQAP-VLA: A Synergistic Quantization-Aware Pruning Framework for High-Performance Vision-Language-Action Models
by: Fang, Hengyu, et al.
Published: (2025)
by: Fang, Hengyu, et al.
Published: (2025)
Multi-identity Human Image Animation with Structural Video Diffusion
by: Wang, Zhenzhi, et al.
Published: (2025)
by: Wang, Zhenzhi, et al.
Published: (2025)
Towards Realistic Scene Generation with LiDAR Diffusion Models
by: Ran, Haoxi, et al.
Published: (2024)
by: Ran, Haoxi, et al.
Published: (2024)
FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning
by: Cao, Jiajun, et al.
Published: (2025)
by: Cao, Jiajun, et al.
Published: (2025)
SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
by: Wang, Hanzhen, et al.
Published: (2025)
by: Wang, Hanzhen, et al.
Published: (2025)
Fast and Memory-Efficient Video Diffusion Using Streamlined Inference
by: Zhan, Zheng, et al.
Published: (2024)
by: Zhan, Zheng, et al.
Published: (2024)
Direct Numerical Layout Generation for 3D Indoor Scene Synthesis via Spatial Reasoning
by: Ran, Xingjian, et al.
Published: (2025)
by: Ran, Xingjian, et al.
Published: (2025)
Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model
by: Yang, Yang, et al.
Published: (2025)
by: Yang, Yang, et al.
Published: (2025)
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
by: Li, Zongjian, et al.
Published: (2024)
by: Li, Zongjian, et al.
Published: (2024)
StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
by: Jin, Xinqi, et al.
Published: (2025)
by: Jin, Xinqi, et al.
Published: (2025)
Context-Aware Temporal Embedding of Objects in Video Data
by: Farhan, Ahnaf, et al.
Published: (2024)
by: Farhan, Ahnaf, et al.
Published: (2024)
Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
by: Liu, Ziyan, et al.
Published: (2025)
by: Liu, Ziyan, et al.
Published: (2025)
TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
by: Huang, Victor Shea-Jay, et al.
Published: (2025)
by: Huang, Victor Shea-Jay, et al.
Published: (2025)
ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation
by: Mai, Ziyang, et al.
Published: (2025)
by: Mai, Ziyang, et al.
Published: (2025)
Dynamic-Aware Video Distillation: Optimizing Temporal Resolution Based on Video Semantics
by: Zhao, Yinjie, et al.
Published: (2025)
by: Zhao, Yinjie, et al.
Published: (2025)
GalaxyDiT: Efficient Video Generation with Guidance Alignment and Adaptive Proxy in Diffusion Transformers
by: Song, Zhiye, et al.
Published: (2025)
by: Song, Zhiye, et al.
Published: (2025)
LD-Pruner: Efficient Pruning of Latent Diffusion Models using Task-Agnostic Insights
by: Castells, Thibault, et al.
Published: (2024)
by: Castells, Thibault, et al.
Published: (2024)
FingER: Content Aware Fine-grained Evaluation with Reasoning for AI-Generated Videos
by: Chen, Rui, et al.
Published: (2025)
by: Chen, Rui, et al.
Published: (2025)
EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation
by: Yang, Songlin, et al.
Published: (2026)
by: Yang, Songlin, et al.
Published: (2026)
TransPrune: Token Transition Pruning for Efficient Large Vision-Language Model
by: Li, Ao, et al.
Published: (2025)
by: Li, Ao, et al.
Published: (2025)
Revealing Temporal Label Noise in Multimodal Hateful Video Classification
by: Yang, Shuonan, et al.
Published: (2025)
by: Yang, Shuonan, et al.
Published: (2025)
Similar Items
-
TokenDance: Token-to-Token Music-to-Dance Generation with Bidirectional Mamba
by: Yang, Ziyue, et al.
Published: (2026) -
EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent
by: Li, Jiaao, et al.
Published: (2025) -
Manifold-Aware Exploration for Reinforcement Learning in Video Generation
by: Zheng, Mingzhe, et al.
Published: (2026) -
HierarchicalPrune: Position-Aware Compression for Large-Scale Diffusion Models
by: Kwon, Young D., et al.
Published: (2025) -
Cluster-based Video Summarization with Temporal Context Awareness
by: Huynh-Lam, Hai-Dang, et al.
Published: (2024)