MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Haozhe, Liu, Shikun, Zhou, Zijian, Xu, Mengmeng, Xie, Yanping, Han, Xiao, Pérez, Juan C., Liu, Ding, Kahatapitiya, Kumara, Jia, Menglin, Wu, Jui-Chieh, He, Sen, Xiang, Tao, Schmidhuber, Jürgen, Pérez-Rúa, Juan-Manuel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Caching for Faster Video Generation with Diffusion Transformers
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
Faster Diffusion via Temporal Attention Decomposition
by: Liu, Haozhe, et al.
Published: (2024)
by: Liu, Haozhe, et al.
Published: (2024)
Understanding Long Videos with Multimodal Language Models
by: Ranasinghe, Kanchana, et al.
Published: (2024)
by: Ranasinghe, Kanchana, et al.
Published: (2024)
GenTron: Diffusion Transformers for Image and Video Generation
by: Chen, Shoufa, et al.
Published: (2023)
by: Chen, Shoufa, et al.
Published: (2023)
OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
by: An, Zhaochong, et al.
Published: (2025)
by: An, Zhaochong, et al.
Published: (2025)
Language Repository for Long Video Understanding
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
by: Liu, Haozhe, et al.
Published: (2025)
by: Liu, Haozhe, et al.
Published: (2025)
VicTR: Video-conditioned Text Representations for Activity Recognition
by: Kahatapitiya, Kumara, et al.
Published: (2023)
by: Kahatapitiya, Kumara, et al.
Published: (2023)
Learning Flow Fields in Attention for Controllable Person Image Generation
by: Zhou, Zijian, et al.
Published: (2024)
by: Zhou, Zijian, et al.
Published: (2024)
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
by: Qiu, Haonan, et al.
Published: (2025)
by: Qiu, Haonan, et al.
Published: (2025)
Object-Centric Diffusion for Efficient Video Editing
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
Scaling Sequence-to-Sequence Generative Neural Rendering
by: Liu, Shikun, et al.
Published: (2025)
by: Liu, Shikun, et al.
Published: (2025)
Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories
by: Jang, Wonbong, et al.
Published: (2026)
by: Jang, Wonbong, et al.
Published: (2026)
Move Anything with Layered Scene Diffusion
by: Ren, Jiawei, et al.
Published: (2024)
by: Ren, Jiawei, et al.
Published: (2024)
Hyper-VolTran: Fast and Generalizable One-Shot Image to 3D Object Structure via HyperNetworks
by: Simon, Christian, et al.
Published: (2023)
by: Simon, Christian, et al.
Published: (2023)
Can Video Diffusion Model Reconstruct 4D Geometry?
by: Mai, Jinjie, et al.
Published: (2025)
by: Mai, Jinjie, et al.
Published: (2025)
Autoregressive Image Generation with Masked Bit Modeling
by: Yu, Qihang, et al.
Published: (2026)
by: Yu, Qihang, et al.
Published: (2026)
Lazy Layers to Make Fine-Tuned Diffusion Models More Traceable
by: Liu, Haozhe, et al.
Published: (2024)
by: Liu, Haozhe, et al.
Published: (2024)
Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA
by: Park, Jongwoo, et al.
Published: (2024)
by: Park, Jongwoo, et al.
Published: (2024)
FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
by: Cong, Yuren, et al.
Published: (2023)
by: Cong, Yuren, et al.
Published: (2023)
Progressive Autoregressive Video Diffusion Models
by: Xie, Desai, et al.
Published: (2024)
by: Xie, Desai, et al.
Published: (2024)
Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression
by: Wang, Lirui, et al.
Published: (2025)
by: Wang, Lirui, et al.
Published: (2025)
Scaling Zero-Shot Reference-to-Video Generation
by: Zhou, Zijian, et al.
Published: (2025)
by: Zhou, Zijian, et al.
Published: (2025)
Towards Robust and Controllable Text-to-Motion via Masked Autoregressive Diffusion
by: Zhang, Zongye, et al.
Published: (2025)
by: Zhang, Zongye, et al.
Published: (2025)
TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
by: Liu, Zhiheng, et al.
Published: (2025)
by: Liu, Zhiheng, et al.
Published: (2025)
Cakrawala Dini
Published: (2022)
Published: (2022)
Dini Araştırmalar
Published: (2018)
Published: (2018)
Distinguished Quantized Guidance for Diffusion-based Sequence Recommendation
by: Mao, Wenyu, et al.
Published: (2025)
by: Mao, Wenyu, et al.
Published: (2025)
Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
by: Liu, Kunhao, et al.
Published: (2025)
by: Liu, Kunhao, et al.
Published: (2025)
La generación femenina de 1950 y el cambio social (1950-2000)
by: Manuel Pérez Rúa
Published: (2013)
by: Manuel Pérez Rúa
Published: (2013)
Taming Teacher Forcing for Masked Autoregressive Video Generation
by: Zhou, Deyu, et al.
Published: (2025)
by: Zhou, Deyu, et al.
Published: (2025)
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling
by: Ye, Zhen, et al.
Published: (2026)
by: Ye, Zhen, et al.
Published: (2026)
EndoGen: Conditional Autoregressive Endoscopic Video Generation
by: Liu, Xinyu, et al.
Published: (2025)
by: Liu, Xinyu, et al.
Published: (2025)
İdrak Dini Araştırmalar Dergisi
Published: (2025)
Published: (2025)
Umde Dini Tetkikler Dergisi
Published: (2022)
Published: (2022)
Marife Dini Araştırmalar Dergisi
Published: (2019)
Published: (2019)
MARRS: Masked Autoregressive Unit-based Reaction Synthesis
by: Wang, Yabiao, et al.
Published: (2025)
by: Wang, Yabiao, et al.
Published: (2025)
One-Forcing: Towards Stable One-Step Autoregressive Video Generation
by: Feng, Jiaqi, et al.
Published: (2026)
by: Feng, Jiaqi, et al.
Published: (2026)
CanvasMAR: Improving Masked Autoregressive Video Prediction With Canvas
by: Li, Zian, et al.
Published: (2025)
by: Li, Zian, et al.
Published: (2025)
Streaming Autoregressive Video Generation via Diagonal Distillation
by: Liu, Jinxiu, et al.
Published: (2026)
by: Liu, Jinxiu, et al.
Published: (2026)
Similar Items
-
Adaptive Caching for Faster Video Generation with Diffusion Transformers
by: Kahatapitiya, Kumara, et al.
Published: (2024) -
Faster Diffusion via Temporal Attention Decomposition
by: Liu, Haozhe, et al.
Published: (2024) -
Understanding Long Videos with Multimodal Language Models
by: Ranasinghe, Kanchana, et al.
Published: (2024) -
GenTron: Diffusion Transformers for Image and Video Generation
by: Chen, Shoufa, et al.
Published: (2023) -
OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
by: An, Zhaochong, et al.
Published: (2025)