Pretraining Frame Preservation for Lightweight Autoregressive Video History Embedding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Lvmin, Cai, Shengqu, Li, Muyang, Zeng, Chong, Lu, Beijia, Rao, Anyi, Han, Song, Wetzstein, Gordon, Agrawala, Maneesh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911500157845504
author Zhang, Lvmin
Cai, Shengqu
Li, Muyang
Zeng, Chong
Lu, Beijia
Rao, Anyi
Han, Song
Wetzstein, Gordon
Agrawala, Maneesh
author_facet Zhang, Lvmin
Cai, Shengqu
Li, Muyang
Zeng, Chong
Lu, Beijia
Rao, Anyi
Han, Song
Wetzstein, Gordon
Agrawala, Maneesh
contents Autoregressive video generation relies on history context for content consistency and storytelling. As video histories grow longer, efficiently encoding them remains an open problem - particularly for personal users and local workflows where compute and memory budgets are limited. We present a lightweight history encoder that maps long video histories into short-length embeddings, pretrained with a frame query objective that learns to attend to content features at arbitrary temporal positions. The pretraining stage provides the encoder with dense history coverage on large-scale video data; the subsequent finetuning stage adapts the pretrained encoder under an autoregressive video generation objective to establish content-level consistency. In this way, the lightweight embeddings achieve comparable performance to heavier alternatives. We evaluate the framework with ablative settings and discuss the architecture designs.
format Preprint
id arxiv_https___arxiv_org_abs_2512_23851
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pretraining Frame Preservation for Lightweight Autoregressive Video History Embedding
Zhang, Lvmin
Cai, Shengqu
Li, Muyang
Zeng, Chong
Lu, Beijia
Rao, Anyi
Han, Song
Wetzstein, Gordon
Agrawala, Maneesh
Computer Vision and Pattern Recognition
Autoregressive video generation relies on history context for content consistency and storytelling. As video histories grow longer, efficiently encoding them remains an open problem - particularly for personal users and local workflows where compute and memory budgets are limited. We present a lightweight history encoder that maps long video histories into short-length embeddings, pretrained with a frame query objective that learns to attend to content features at arbitrary temporal positions. The pretraining stage provides the encoder with dense history coverage on large-scale video data; the subsequent finetuning stage adapts the pretrained encoder under an autoregressive video generation objective to establish content-level consistency. In this way, the lightweight embeddings achieve comparable performance to heavier alternatives. We evaluate the framework with ablative settings and discuss the architecture designs.
title Pretraining Frame Preservation for Lightweight Autoregressive Video History Embedding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.23851