Taming Teacher Forcing for Masked Autoregressive Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Deyu, Sun, Quan, Peng, Yuang, Yan, Kun, Dong, Runpei, Wang, Duomin, Ge, Zheng, Duan, Nan, Zhang, Xiangyu, Ni, Lionel M., Shum, Heung-Yeung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912197748195328
author Zhou, Deyu
Sun, Quan
Peng, Yuang
Yan, Kun
Dong, Runpei
Wang, Duomin
Ge, Zheng
Duan, Nan
Zhang, Xiangyu
Ni, Lionel M.
Shum, Heung-Yeung
author_facet Zhou, Deyu
Sun, Quan
Peng, Yuang
Yan, Kun
Dong, Runpei
Wang, Duomin
Ge, Zheng
Duan, Nan
Zhang, Xiangyu
Ni, Lionel M.
Shum, Heung-Yeung
contents We introduce MAGI, a hybrid video generation framework that combines masked modeling for intra-frame generation with causal modeling for next-frame generation. Our key innovation, Complete Teacher Forcing (CTF), conditions masked frames on complete observation frames rather than masked ones (namely Masked Teacher Forcing, MTF), enabling a smooth transition from token-level (patch-level) to frame-level autoregressive generation. CTF significantly outperforms MTF, achieving a +23% improvement in FVD scores on first-frame conditioned video prediction. To address issues like exposure bias, we employ targeted training strategies, setting a new benchmark in autoregressive video generation. Experiments show that MAGI can generate long, coherent video sequences exceeding 100 frames, even when trained on as few as 16 frames, highlighting its potential for scalable, high-quality video generation.
format Preprint
id arxiv_https___arxiv_org_abs_2501_12389
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Taming Teacher Forcing for Masked Autoregressive Video Generation
Zhou, Deyu
Sun, Quan
Peng, Yuang
Yan, Kun
Dong, Runpei
Wang, Duomin
Ge, Zheng
Duan, Nan
Zhang, Xiangyu
Ni, Lionel M.
Shum, Heung-Yeung
Computer Vision and Pattern Recognition
We introduce MAGI, a hybrid video generation framework that combines masked modeling for intra-frame generation with causal modeling for next-frame generation. Our key innovation, Complete Teacher Forcing (CTF), conditions masked frames on complete observation frames rather than masked ones (namely Masked Teacher Forcing, MTF), enabling a smooth transition from token-level (patch-level) to frame-level autoregressive generation. CTF significantly outperforms MTF, achieving a +23% improvement in FVD scores on first-frame conditioned video prediction. To address issues like exposure bias, we employ targeted training strategies, setting a new benchmark in autoregressive video generation. Experiments show that MAGI can generate long, coherent video sequences exceeding 100 frames, even when trained on as few as 16 frames, highlighting its potential for scalable, high-quality video generation.
title Taming Teacher Forcing for Masked Autoregressive Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.12389