Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Lvmin, Cai, Shengqu, Li, Muyang, Wetzstein, Gordon, Agrawala, Maneesh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909846633185280
author Zhang, Lvmin
Cai, Shengqu
Li, Muyang
Wetzstein, Gordon
Agrawala, Maneesh
author_facet Zhang, Lvmin
Cai, Shengqu
Li, Muyang
Wetzstein, Gordon
Agrawala, Maneesh
contents We present a neural network structure, FramePack, to train next-frame (or next-frame-section) prediction models for video generation. FramePack compresses input frame contexts with frame-wise importance so that more frames can be encoded within a fixed context length, with more important frames having longer contexts. The frame importance can be measured using time proximity, feature similarity, or hybrid metrics. The packing method allows for inference with thousands of frames and training with relatively large batch sizes. We also present drift prevention methods to address observation bias (error accumulation), including early-established endpoints, adjusted sampling orders, and discrete history representation. Ablation studies validate the effectiveness of the anti-drifting methods in both single-directional video streaming and bi-directional video generation. Finally, we show that existing video diffusion models can be finetuned with FramePack, and analyze the differences between different packing schedules.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12626
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
Zhang, Lvmin
Cai, Shengqu
Li, Muyang
Wetzstein, Gordon
Agrawala, Maneesh
Computer Vision and Pattern Recognition
We present a neural network structure, FramePack, to train next-frame (or next-frame-section) prediction models for video generation. FramePack compresses input frame contexts with frame-wise importance so that more frames can be encoded within a fixed context length, with more important frames having longer contexts. The frame importance can be measured using time proximity, feature similarity, or hybrid metrics. The packing method allows for inference with thousands of frames and training with relatively large batch sizes. We also present drift prevention methods to address observation bias (error accumulation), including early-established endpoints, adjusted sampling orders, and discrete history representation. Ablation studies validate the effectiveness of the anti-drifting methods in both single-directional video streaming and bi-directional video generation. Finally, we show that existing video diffusion models can be finetuned with FramePack, and analyze the differences between different packing schedules.
title Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.12626