Mixture of Contexts for Long Video Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cai, Shengqu, Yang, Ceyuan, Zhang, Lvmin, Guo, Yuwei, Xiao, Junfei, Yang, Ziyan, Xu, Yinghao, Yang, Zhenheng, Yuille, Alan, Guibas, Leonidas, Agrawala, Maneesh, Jiang, Lu, Wetzstein, Gordon
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918239484772352
author Cai, Shengqu
Yang, Ceyuan
Zhang, Lvmin
Guo, Yuwei
Xiao, Junfei
Yang, Ziyan
Xu, Yinghao
Yang, Zhenheng
Yuille, Alan
Guibas, Leonidas
Agrawala, Maneesh
Jiang, Lu
Wetzstein, Gordon
author_facet Cai, Shengqu
Yang, Ceyuan
Zhang, Lvmin
Guo, Yuwei
Xiao, Junfei
Yang, Ziyan
Xu, Yinghao
Yang, Zhenheng
Yuille, Alan
Guibas, Leonidas
Agrawala, Maneesh
Jiang, Lu
Wetzstein, Gordon
contents Long video generation is fundamentally a long context memory problem: models must retain and retrieve salient events across a long range without collapsing or drifting. However, scaling diffusion transformers to generate long-context videos is fundamentally limited by the quadratic cost of self-attention, which makes memory and computation intractable and difficult to optimize for long sequences. We recast long-context video generation as an internal information retrieval task and propose a simple, learnable sparse attention routing module, Mixture of Contexts (MoC), as an effective long-term memory retrieval engine. In MoC, each query dynamically selects a few informative chunks plus mandatory anchors (caption, local windows) to attend to, with causal routing that prevents loop closures. As we scale the data and gradually sparsify the routing, the model allocates compute to salient history, preserving identities, actions, and scenes over minutes of content. Efficiency follows as a byproduct of retrieval (near-linear scaling), which enables practical training and synthesis, and the emergence of memory and consistency at the scale of minutes.
format Preprint
id arxiv_https___arxiv_org_abs_2508_21058
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mixture of Contexts for Long Video Generation
Cai, Shengqu
Yang, Ceyuan
Zhang, Lvmin
Guo, Yuwei
Xiao, Junfei
Yang, Ziyan
Xu, Yinghao
Yang, Zhenheng
Yuille, Alan
Guibas, Leonidas
Agrawala, Maneesh
Jiang, Lu
Wetzstein, Gordon
Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
Long video generation is fundamentally a long context memory problem: models must retain and retrieve salient events across a long range without collapsing or drifting. However, scaling diffusion transformers to generate long-context videos is fundamentally limited by the quadratic cost of self-attention, which makes memory and computation intractable and difficult to optimize for long sequences. We recast long-context video generation as an internal information retrieval task and propose a simple, learnable sparse attention routing module, Mixture of Contexts (MoC), as an effective long-term memory retrieval engine. In MoC, each query dynamically selects a few informative chunks plus mandatory anchors (caption, local windows) to attend to, with causal routing that prevents loop closures. As we scale the data and gradually sparsify the routing, the model allocates compute to salient history, preserving identities, actions, and scenes over minutes of content. Efficiency follows as a byproduct of retrieval (near-linear scaling), which enables practical training and synthesis, and the emergence of memory and consistency at the scale of minutes.
title Mixture of Contexts for Long Video Generation
topic Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.21058