PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Korzhenkov, Denis, Karjauv, Adil, Karnewar, Animesh, Ghafoorian, Mohsen, Habibian, Amirhossein |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
by: Ghafoorian, Mohsen, et al.
Published: (2025)
by: Ghafoorian, Mohsen, et al.
Published: (2025)
Neodragon: Mobile Video Generation using Diffusion Transformer
by: Karnewar, Animesh, et al.
Published: (2025)
by: Karnewar, Animesh, et al.
Published: (2025)
ReHyAt: Recurrent Hybrid Attention for Video Diffusion Transformers
by: Ghafoorian, Mohsen, et al.
Published: (2026)
by: Ghafoorian, Mohsen, et al.
Published: (2026)
MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
by: Bhowmik, Aritra, et al.
Published: (2025)
by: Bhowmik, Aritra, et al.
Published: (2025)
MoViE: Mobile Diffusion for Video Editing
by: Karjauv, Adil, et al.
Published: (2024)
by: Karjauv, Adil, et al.
Published: (2024)
Object-Centric Diffusion for Efficient Video Editing
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
Mobile Video Diffusion
by: Yahia, Haitam Ben, et al.
Published: (2024)
by: Yahia, Haitam Ben, et al.
Published: (2024)
TPDiff: Temporal Pyramid Video Diffusion Model
by: Ran, Lingmin, et al.
Published: (2025)
by: Ran, Lingmin, et al.
Published: (2025)
Pyramidal Flow Matching for Efficient Video Generative Modeling
by: Jin, Yang, et al.
Published: (2024)
by: Jin, Yang, et al.
Published: (2024)
Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation
by: Chen, Jiayu, et al.
Published: (2026)
by: Chen, Jiayu, et al.
Published: (2026)
Dynamic Pyramid Network for Efficient Multimodal Large Language Model
by: Ai, Hao, et al.
Published: (2025)
by: Ai, Hao, et al.
Published: (2025)
Face Pyramid Vision Transformer
by: Islam, Khawar, et al.
Published: (2022)
by: Islam, Khawar, et al.
Published: (2022)
PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
by: Xing, Long, et al.
Published: (2024)
by: Xing, Long, et al.
Published: (2024)
PyramidMamba: Rethinking Pyramid Feature Fusion with Selective Space State Model for Semantic Segmentation of Remote Sensing Imagery
by: Wang, Libo, et al.
Published: (2024)
by: Wang, Libo, et al.
Published: (2024)
FastCAD: Real-Time CAD Retrieval and Alignment from Scans and Videos
by: Langer, Florian, et al.
Published: (2024)
by: Langer, Florian, et al.
Published: (2024)
Efficient Pyramid Channel Attention Network for Pathological Myopia Recognition
by: Zhang, Xiaoqing, et al.
Published: (2023)
by: Zhang, Xiaoqing, et al.
Published: (2023)
Multi-Scale Local Speculative Decoding for Image Generation
by: Peruzzo, Elia, et al.
Published: (2026)
by: Peruzzo, Elia, et al.
Published: (2026)
HTD-Mamba: Efficient Hyperspectral Target Detection with Pyramid State Space Model
by: Shen, Dunbin, et al.
Published: (2024)
by: Shen, Dunbin, et al.
Published: (2024)
Mumpy: Multilateral Temporal-view Pyramid Transformer for Video Inpainting Detection
by: Zhang, Ying, et al.
Published: (2024)
by: Zhang, Ying, et al.
Published: (2024)
PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation
by: Li, Xiaolong, et al.
Published: (2025)
by: Li, Xiaolong, et al.
Published: (2025)
Pyramidal Patchification Flow for Visual Generation
by: Li, Hui, et al.
Published: (2025)
by: Li, Hui, et al.
Published: (2025)
Parameter-Inverted Image Pyramid Networks
by: Zhu, Xizhou, et al.
Published: (2024)
by: Zhu, Xizhou, et al.
Published: (2024)
Scene-Aware Location Modeling for Data Augmentation in Automotive Object Detection
by: Petersen, Jens, et al.
Published: (2025)
by: Petersen, Jens, et al.
Published: (2025)
PNeRV: Enhancing Spatial Consistency via Pyramidal Neural Representation for Videos
by: Zhao, Qi, et al.
Published: (2024)
by: Zhao, Qi, et al.
Published: (2024)
Real-Time Video Generation with Pyramid Attention Broadcast
by: Zhao, Xuanlei, et al.
Published: (2024)
by: Zhao, Xuanlei, et al.
Published: (2024)
Enhancing Novel View Synthesis via Geometry Grounded Set Diffusion
by: Zanjani, Farhad G., et al.
Published: (2026)
by: Zanjani, Farhad G., et al.
Published: (2026)
Ultra-High-Resolution Image Synthesis with Pyramid Diffusion Model
by: Yang, Jiajie
Published: (2024)
by: Yang, Jiajie
Published: (2024)
Pyramid Attention Network for Medical Image Registration
by: Wang, Zhuoyuan, et al.
Published: (2024)
by: Wang, Zhuoyuan, et al.
Published: (2024)
Pyramidal Adaptive Cross-Gating for Multimodal Detection
by: Gu, Zidong, et al.
Published: (2025)
by: Gu, Zidong, et al.
Published: (2025)
Pyramid Hierarchical Transformer for Hyperspectral Image Classification
by: Ahmad, Muhammad, et al.
Published: (2024)
by: Ahmad, Muhammad, et al.
Published: (2024)
Global Feature Pyramid Network
by: Xiao, Weilin, et al.
Published: (2023)
by: Xiao, Weilin, et al.
Published: (2023)
PyramidStyler: Transformer-Based Neural Style Transfer with Pyramidal Positional Encoding and Reinforcement Learning
by: Durairaju, Raahul Krishna, et al.
Published: (2025)
by: Durairaju, Raahul Krishna, et al.
Published: (2025)
UniVid: Pyramid Diffusion Model for High Quality Video Generation
by: Xiao, Xinyu, et al.
Published: (2026)
by: Xiao, Xinyu, et al.
Published: (2026)
LaFAM: Unsupervised Feature Attribution with Label-free Activation Maps
by: Karjauv, Aray, et al.
Published: (2024)
by: Karjauv, Aray, et al.
Published: (2024)
Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
by: Chen, Yanlong, et al.
Published: (2026)
by: Chen, Yanlong, et al.
Published: (2026)
Learning Inverse Laplacian Pyramid for Progressive Depth Completion
by: Wang, Kun, et al.
Published: (2025)
by: Wang, Kun, et al.
Published: (2025)
FIPGNet:Pyramid grafting network with feature interaction strategies
by: Ding, Ziyi, et al.
Published: (2024)
by: Ding, Ziyi, et al.
Published: (2024)
Rethinking Features-Fused-Pyramid-Neck for Object Detection
by: Li, Hulin
Published: (2025)
by: Li, Hulin
Published: (2025)
Pyramid Feature Attention Network for Monocular Depth Prediction
by: Xu, Yifang, et al.
Published: (2024)
by: Xu, Yifang, et al.
Published: (2024)
PyraTok: Language-Aligned Pyramidal Tokenizer for Video Understanding and Generation
by: Susladkar, Onkar, et al.
Published: (2026)
by: Susladkar, Onkar, et al.
Published: (2026)
Similar Items
-
Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
by: Ghafoorian, Mohsen, et al.
Published: (2025) -
Neodragon: Mobile Video Generation using Diffusion Transformer
by: Karnewar, Animesh, et al.
Published: (2025) -
ReHyAt: Recurrent Hybrid Attention for Video Diffusion Transformers
by: Ghafoorian, Mohsen, et al.
Published: (2026) -
MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
by: Bhowmik, Aritra, et al.
Published: (2025) -
MoViE: Mobile Diffusion for Video Editing
by: Karjauv, Adil, et al.
Published: (2024)