VMonarch: Efficient Video Diffusion Transformers with Structured Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Cheng, Chen, Haoxian, Hou, Liang, Fan, Qi, Wu, Gangshan, Tao, Xin, Wang, Limin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DDT: Decoupled Diffusion Transformer
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers
by: Chen, Pengtao, et al.
Published: (2025)
by: Chen, Pengtao, et al.
Published: (2025)
SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
by: Chen, Junsong, et al.
Published: (2025)
by: Chen, Junsong, et al.
Published: (2025)
Re-Attentional Controllable Video Diffusion Editing
by: Wang, Yuanzhi, et al.
Published: (2024)
by: Wang, Yuanzhi, et al.
Published: (2024)
SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
MixFormerV2: Efficient Fully Transformer Tracking
by: Cui, Yutao, et al.
Published: (2023)
by: Cui, Yutao, et al.
Published: (2023)
LiteAttention: A Temporal Sparse Attention for Diffusion Transformers
by: Shmilovich, Dor, et al.
Published: (2025)
by: Shmilovich, Dor, et al.
Published: (2025)
RASA: Replace Anyone, Say Anything -- A Training-Free Framework for Audio-Driven and Universal Portrait Video Editing
by: Pan, Tianrui, et al.
Published: (2025)
by: Pan, Tianrui, et al.
Published: (2025)
USV: Towards Understanding the User-generated Short-form Videos
by: Cheng, Haoyue, et al.
Published: (2026)
by: Cheng, Haoyue, et al.
Published: (2026)
LuxDiT: Lighting Estimation with Video Diffusion Transformer
by: Liang, Ruofan, et al.
Published: (2025)
by: Liang, Ruofan, et al.
Published: (2025)
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
by: Feng, Wenfeng, et al.
Published: (2025)
by: Feng, Wenfeng, et al.
Published: (2025)
Make It Efficient: Dynamic Sparse Attention for Autoregressive Image Generation
by: Xiang, Xunzhi, et al.
Published: (2025)
by: Xiang, Xunzhi, et al.
Published: (2025)
From Structure to Detail: Hierarchical Distillation for Efficient Diffusion Model
by: Cheng, Hanbo, et al.
Published: (2025)
by: Cheng, Hanbo, et al.
Published: (2025)
RectifiedHR: Enable Efficient High-Resolution Synthesis via Energy Rectification
by: Yang, Zhen, et al.
Published: (2025)
by: Yang, Zhen, et al.
Published: (2025)
DraftAttention: Fast Video Diffusion via Low-Resolution Attention Guidance
by: Shen, Xuan, et al.
Published: (2025)
by: Shen, Xuan, et al.
Published: (2025)
Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance
by: Wu, Song, et al.
Published: (2026)
by: Wu, Song, et al.
Published: (2026)
BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models
by: Shi, Fengyuan, et al.
Published: (2023)
by: Shi, Fengyuan, et al.
Published: (2023)
StarVid: Enhancing Semantic Alignment in Video Diffusion Models via Spatial and SynTactic Guided Attention Refocusing
by: Li, Yuanhang, et al.
Published: (2024)
by: Li, Yuanhang, et al.
Published: (2024)
Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
by: Yan, Xin, et al.
Published: (2024)
by: Yan, Xin, et al.
Published: (2024)
Taming Outlier Tokens in Diffusion Transformers
by: Wu, Xiaoyu, et al.
Published: (2026)
by: Wu, Xiaoyu, et al.
Published: (2026)
Diffusion Attention Expert Model for Predicting and Semi-automatic Localizing STAS in Lung Cancer Histopathological Images
by: Pan, Liangrui, et al.
Published: (2026)
by: Pan, Liangrui, et al.
Published: (2026)
On-Device Diffusion Transformer Policy for Efficient Robot Manipulation
by: Wu, Yiming, et al.
Published: (2025)
by: Wu, Yiming, et al.
Published: (2025)
Probing Routing-Conditional Calibration in Attention-Residual Transformers
by: Liang, Wenhao, et al.
Published: (2026)
by: Liang, Wenhao, et al.
Published: (2026)
DiTFastAttnV2: Head-wise Attention Compression for Multi-Modality Diffusion Transformers
by: Zhang, Hanling, et al.
Published: (2025)
by: Zhang, Hanling, et al.
Published: (2025)
Efficient Video Diffusion with Sparse Information Transmission for Video Compression
by: Zhou, Mingde, et al.
Published: (2026)
by: Zhou, Mingde, et al.
Published: (2026)
Attention Retention for Continual Learning with Vision Transformers
by: Lu, Yue, et al.
Published: (2026)
by: Lu, Yue, et al.
Published: (2026)
Spiking Vision Transformer with Saccadic Attention
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
DAWN: Dynamic Frame Avatar with Non-autoregressive Diffusion Framework for Talking Head Video Generation
by: Cheng, Hanbo, et al.
Published: (2024)
by: Cheng, Hanbo, et al.
Published: (2024)
ASR: Attention-alike Structural Re-parameterization
by: Zhong, Shanshan, et al.
Published: (2023)
by: Zhong, Shanshan, et al.
Published: (2023)
Dual Prompting Image Restoration with Diffusion Transformers
by: Kong, Dehong, et al.
Published: (2025)
by: Kong, Dehong, et al.
Published: (2025)
Towards Precise Scaling Laws for Video Diffusion Transformers
by: Yin, Yuanyang, et al.
Published: (2024)
by: Yin, Yuanyang, et al.
Published: (2024)
SDAR-VL: Stable and Efficient Block-wise Diffusion for Vision-Language Understanding
by: Cheng, Shuang, et al.
Published: (2025)
by: Cheng, Shuang, et al.
Published: (2025)
A Generic Shared Attention Mechanism for Various Backbone Neural Networks
by: Huang, Zhongzhan, et al.
Published: (2022)
by: Huang, Zhongzhan, et al.
Published: (2022)
FAME: Fairness-aware Attention-modulated Video Editing
by: Wu, Zhangkai, et al.
Published: (2025)
by: Wu, Zhangkai, et al.
Published: (2025)
SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces
by: Oshima, Yuta, et al.
Published: (2024)
by: Oshima, Yuta, et al.
Published: (2024)
BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation
by: Wang, Ruotong, et al.
Published: (2025)
by: Wang, Ruotong, et al.
Published: (2025)
IC-Effect: Precise and Efficient Video Effects Editing via In-Context Learning
by: Li, Yuanhang, et al.
Published: (2025)
by: Li, Yuanhang, et al.
Published: (2025)
Exploiting Spatiotemporal Properties for Efficient Event-Driven Human Pose Estimation
by: Zhou, Haoxian, et al.
Published: (2025)
by: Zhou, Haoxian, et al.
Published: (2025)
Fast and Memory-Efficient Video Diffusion Using Streamlined Inference
by: Zhan, Zheng, et al.
Published: (2024)
by: Zhan, Zheng, et al.
Published: (2024)
StreamingTOM: Streaming Token Compression for Efficient Video Understanding
by: Chen, Xueyi, et al.
Published: (2025)
by: Chen, Xueyi, et al.
Published: (2025)
Similar Items
-
DDT: Decoupled Diffusion Transformer
by: Wang, Shuai, et al.
Published: (2025) -
Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers
by: Chen, Pengtao, et al.
Published: (2025) -
SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
by: Chen, Junsong, et al.
Published: (2025) -
Re-Attentional Controllable Video Diffusion Editing
by: Wang, Yuanzhi, et al.
Published: (2024) -
SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
by: Wu, Tao, et al.
Published: (2024)