WorldPack: Compressed Memory Improves Spatial Consistency in Video World Modeling

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Oshima, Yuta, Iwasawa, Yusuke, Suzuki, Masahiro, Matsuo, Yutaka, Furuta, Hiroki
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917118584291328
author Oshima, Yuta
Iwasawa, Yusuke
Suzuki, Masahiro
Matsuo, Yutaka
Furuta, Hiroki
author_facet Oshima, Yuta
Iwasawa, Yusuke
Suzuki, Masahiro
Matsuo, Yutaka
Furuta, Hiroki
contents Video world models have attracted significant attention for their ability to produce high-fidelity future visual observations conditioned on past observations and navigation actions. Temporally- and spatially-consistent, long-term world modeling has been a long-standing problem, unresolved with even recent state-of-the-art models, due to the prohibitively expensive computational costs for long-context inputs. In this paper, we propose WorldPack, a video world model with efficient compressed memory, which significantly improves spatial consistency, fidelity, and quality in long-term generation despite much shorter context length. Our compressed memory consists of trajectory packing and memory retrieval; trajectory packing realizes high context efficiency, and memory retrieval maintains the consistency in rollouts and helps long-term generations that require spatial reasoning. Our performance is evaluated with LoopNav, a benchmark on Minecraft, specialized for the evaluation of long-term consistency, and we verify that WorldPack notably outperforms strong state-of-the-art models.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02473
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WorldPack: Compressed Memory Improves Spatial Consistency in Video World Modeling
Oshima, Yuta
Iwasawa, Yusuke
Suzuki, Masahiro
Matsuo, Yutaka
Furuta, Hiroki
Computer Vision and Pattern Recognition
Machine Learning
Video world models have attracted significant attention for their ability to produce high-fidelity future visual observations conditioned on past observations and navigation actions. Temporally- and spatially-consistent, long-term world modeling has been a long-standing problem, unresolved with even recent state-of-the-art models, due to the prohibitively expensive computational costs for long-context inputs. In this paper, we propose WorldPack, a video world model with efficient compressed memory, which significantly improves spatial consistency, fidelity, and quality in long-term generation despite much shorter context length. Our compressed memory consists of trajectory packing and memory retrieval; trajectory packing realizes high context efficiency, and memory retrieval maintains the consistency in rollouts and helps long-term generations that require spatial reasoning. Our performance is evaluated with LoopNav, a benchmark on Minecraft, specialized for the evaluation of long-term consistency, and we verify that WorldPack notably outperforms strong state-of-the-art models.
title WorldPack: Compressed Memory Improves Spatial Consistency in Video World Modeling
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2512.02473