LongVie 2: Multimodal Controllable Ultra-Long Video World Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Jianxiong, Chen, Zhaoxi, Liu, Xian, Zhuang, Junhao, Xu, Chengming, Feng, Jianfeng, Qiao, Yu, Fu, Yanwei, Si, Chenyang, Liu, Ziwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908713211658240
author Gao, Jianxiong
Chen, Zhaoxi
Liu, Xian
Zhuang, Junhao
Xu, Chengming
Feng, Jianfeng
Qiao, Yu
Fu, Yanwei
Si, Chenyang
Liu, Ziwei
author_facet Gao, Jianxiong
Chen, Zhaoxi
Liu, Xian
Zhuang, Junhao
Xu, Chengming
Feng, Jianfeng
Qiao, Yu
Fu, Yanwei
Si, Chenyang
Liu, Ziwei
contents Building video world models upon pretrained video generation systems represents an important yet challenging step toward general spatiotemporal intelligence. A world model should possess three essential properties: controllability, long-term visual quality, and temporal consistency. To this end, we take a progressive approach-first enhancing controllability and then extending toward long-term, high-quality generation. We present LongVie 2, an end-to-end autoregressive framework trained in three stages: (1) Multi-modal guidance, which integrates dense and sparse control signals to provide implicit world-level supervision and improve controllability; (2) Degradation-aware training on the input frame, bridging the gap between training and long-term inference to maintain high visual quality; and (3) History-context guidance, which aligns contextual information across adjacent clips to ensure temporal consistency. We further introduce LongVGenBench, a comprehensive benchmark comprising 100 high-resolution one-minute videos covering diverse real-world and synthetic environments. Extensive experiments demonstrate that LongVie 2 achieves state-of-the-art performance in long-range controllability, temporal coherence, and visual fidelity, and supports continuous video generation lasting up to five minutes, marking a significant step toward unified video world modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13604
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LongVie 2: Multimodal Controllable Ultra-Long Video World Model
Gao, Jianxiong
Chen, Zhaoxi
Liu, Xian
Zhuang, Junhao
Xu, Chengming
Feng, Jianfeng
Qiao, Yu
Fu, Yanwei
Si, Chenyang
Liu, Ziwei
Computer Vision and Pattern Recognition
Building video world models upon pretrained video generation systems represents an important yet challenging step toward general spatiotemporal intelligence. A world model should possess three essential properties: controllability, long-term visual quality, and temporal consistency. To this end, we take a progressive approach-first enhancing controllability and then extending toward long-term, high-quality generation. We present LongVie 2, an end-to-end autoregressive framework trained in three stages: (1) Multi-modal guidance, which integrates dense and sparse control signals to provide implicit world-level supervision and improve controllability; (2) Degradation-aware training on the input frame, bridging the gap between training and long-term inference to maintain high visual quality; and (3) History-context guidance, which aligns contextual information across adjacent clips to ensure temporal consistency. We further introduce LongVGenBench, a comprehensive benchmark comprising 100 high-resolution one-minute videos covering diverse real-world and synthetic environments. Extensive experiments demonstrate that LongVie 2 achieves state-of-the-art performance in long-range controllability, temporal coherence, and visual fidelity, and supports continuous video generation lasting up to five minutes, marking a significant step toward unified video world modeling.
title LongVie 2: Multimodal Controllable Ultra-Long Video World Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.13604