VideoMAR: Autoregressive Video Generatio with Continuous Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Hu, Gong, Biao, Yuan, Hangjie, Zheng, DanDan, Chai, Weilong, Chen, Jingdong, Zheng, Kecheng, Zhao, Feng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Frequency Autoregressive Image Generation with Continuous Tokens
by: Yu, Hu, et al.
Published: (2025)
by: Yu, Hu, et al.
Published: (2025)
SpeedUpNet: A Plug-and-Play Adapter Network for Accelerating Text-to-Image Diffusion Models
by: Chai, Weilong, et al.
Published: (2023)
by: Chai, Weilong, et al.
Published: (2023)
Mimir: Improving Video Diffusion Models for Precise Text Understanding
by: Tan, Shuai, et al.
Published: (2024)
by: Tan, Shuai, et al.
Published: (2024)
Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
by: Huang, Ziyuan, et al.
Published: (2025)
by: Huang, Ziyuan, et al.
Published: (2025)
VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations
by: Qiao, Qianqian, et al.
Published: (2025)
by: Qiao, Qianqian, et al.
Published: (2025)
VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation
by: Jiang, Longteng, et al.
Published: (2026)
by: Jiang, Longteng, et al.
Published: (2026)
MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation
by: Shi, Shuwei, et al.
Published: (2024)
by: Shi, Shuwei, et al.
Published: (2024)
Animate-X: Universal Character Image Animation with Enhanced Motion Representation
by: Tan, Shuai, et al.
Published: (2024)
by: Tan, Shuai, et al.
Published: (2024)
LumiSculpt: Enabling Consistent Portrait Lighting in Video Generation
by: Zhang, Yuxin, et al.
Published: (2024)
by: Zhang, Yuxin, et al.
Published: (2024)
Adaptive Begin-of-Video Tokens for Autoregressive Video Diffusion Models
by: Cheng, Tianle, et al.
Published: (2025)
by: Cheng, Tianle, et al.
Published: (2025)
SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation
by: Tan, Shuai, et al.
Published: (2025)
by: Tan, Shuai, et al.
Published: (2025)
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
by: Xiong, Tianwei, et al.
Published: (2026)
by: Xiong, Tianwei, et al.
Published: (2026)
Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction
by: Wei, Yujie, et al.
Published: (2026)
by: Wei, Yujie, et al.
Published: (2026)
DreamRelation: Relation-Centric Video Customization
by: Wei, Yujie, et al.
Published: (2025)
by: Wei, Yujie, et al.
Published: (2025)
AtomoVideo: High Fidelity Image-to-Video Generation
by: Gong, Litong, et al.
Published: (2024)
by: Gong, Litong, et al.
Published: (2024)
StyleTokenizer: Defining Image Style by a Single Instance for Controlling Diffusion Models
by: Li, Wen, et al.
Published: (2024)
by: Li, Wen, et al.
Published: (2024)
Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight
by: Huang, Ziyuan, et al.
Published: (2024)
by: Huang, Ziyuan, et al.
Published: (2024)
Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
by: Zheng, Peng, et al.
Published: (2025)
by: Zheng, Peng, et al.
Published: (2025)
VideoOrion: Tokenizing Object Dynamics in Videos
by: Feng, Yicheng, et al.
Published: (2024)
by: Feng, Yicheng, et al.
Published: (2024)
Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
by: Liu, Xiangrui, et al.
Published: (2025)
by: Liu, Xiangrui, et al.
Published: (2025)
Long-Context Autoregressive Video Modeling with Next-Frame Prediction
by: Gu, Yuchao, et al.
Published: (2025)
by: Gu, Yuchao, et al.
Published: (2025)
Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
by: Yuan, Hangjie, et al.
Published: (2025)
by: Yuan, Hangjie, et al.
Published: (2025)
FastSTAR: Spatiotemporal Token Pruning for Efficient Autoregressive Video Synthesis
by: Yune, Sungwoong, et al.
Published: (2026)
by: Yune, Sungwoong, et al.
Published: (2026)
DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control
by: Wei, Yujie, et al.
Published: (2024)
by: Wei, Yujie, et al.
Published: (2024)
ZeroFlow: Overcoming Catastrophic Forgetting is Easier than You Think
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory
by: Yu, Yifei, et al.
Published: (2025)
by: Yu, Yifei, et al.
Published: (2025)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
CanvasMAR: Improving Masked Autoregressive Video Prediction With Canvas
by: Li, Zian, et al.
Published: (2025)
by: Li, Zian, et al.
Published: (2025)
Multi-Task Label Discovery via Hierarchical Task Tokens for Partially Annotated Dense Predictions
by: Zhang, Jingdong, et al.
Published: (2024)
by: Zhang, Jingdong, et al.
Published: (2024)
HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising
by: Zou, Kai, et al.
Published: (2026)
by: Zou, Kai, et al.
Published: (2026)
TokenTrim: Inference-Time Token Pruning for Autoregressive Long Video Generation
by: Shaulov, Ariel, et al.
Published: (2026)
by: Shaulov, Ariel, et al.
Published: (2026)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
by: Cai, Lingling, et al.
Published: (2024)
by: Cai, Lingling, et al.
Published: (2024)
LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior
by: Wang, Hanyu, et al.
Published: (2024)
by: Wang, Hanyu, et al.
Published: (2024)
Real-Time Motion-Controllable Autoregressive Video Diffusion
by: Zhao, Kesen, et al.
Published: (2025)
by: Zhao, Kesen, et al.
Published: (2025)
Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors
by: Wang, Xiangchen, et al.
Published: (2025)
by: Wang, Xiangchen, et al.
Published: (2025)
DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing
by: Cai, Lingling, et al.
Published: (2025)
by: Cai, Lingling, et al.
Published: (2025)
Video In-context Learning: Autoregressive Transformers are Zero-Shot Video Imitators
by: Zhang, Wentao, et al.
Published: (2024)
by: Zhang, Wentao, et al.
Published: (2024)
Motion-Aware Caching for Efficient Autoregressive Video Generation
by: Xu, Jing, et al.
Published: (2026)
by: Xu, Jing, et al.
Published: (2026)
Tuning-Free Noise Rectification for High Fidelity Image-to-Video Generation
by: Li, Weijie, et al.
Published: (2024)
by: Li, Weijie, et al.
Published: (2024)
UKnow: A Unified Knowledge Protocol with Multimodal Knowledge Graph Datasets for Reasoning and Vision-Language Pre-Training
by: Gong, Biao, et al.
Published: (2023)
by: Gong, Biao, et al.
Published: (2023)
Similar Items
-
Frequency Autoregressive Image Generation with Continuous Tokens
by: Yu, Hu, et al.
Published: (2025) -
SpeedUpNet: A Plug-and-Play Adapter Network for Accelerating Text-to-Image Diffusion Models
by: Chai, Weilong, et al.
Published: (2023) -
Mimir: Improving Video Diffusion Models for Precise Text Understanding
by: Tan, Shuai, et al.
Published: (2024) -
Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
by: Huang, Ziyuan, et al.
Published: (2025) -
VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations
by: Qiao, Qianqian, et al.
Published: (2025)