Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wu, Fan, Wei, Jiacheng, Li, Ruibo, Xu, Yi, Li, Junyou, Ye, Deheng, Lin, Guosheng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2512.02793
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908687711338496
author Wu, Fan
Wei, Jiacheng
Li, Ruibo
Xu, Yi
Li, Junyou
Ye, Deheng
Lin, Guosheng
author_facet Wu, Fan
Wei, Jiacheng
Li, Ruibo
Xu, Yi
Li, Junyou
Ye, Deheng
Lin, Guosheng
contents Video-based world models have recently garnered increasing attention for their ability to synthesize diverse and dynamic visual environments. In this paper, we focus on shared world modeling, where a model generates multiple videos from a set of input images, each representing the same underlying world in different camera poses. We propose IC-World, a novel generation framework, enabling parallel generation for all input images via activating the inherent in-context generation capability of large video models. We further finetune IC-World via reinforcement learning, Group Relative Policy Optimization, together with two proposed novel reward models to enforce scene-level geometry consistency and object-level motion consistency among the set of generated videos. Extensive experiments demonstrate that IC-World substantially outperforms state-of-the-art methods in both geometry and motion consistency. To the best of our knowledge, this is the first work to systematically explore the shared world modeling problem with video-based world models.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02793
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IC-World: In-Context Generation for Shared World Modeling
Wu, Fan
Wei, Jiacheng
Li, Ruibo
Xu, Yi
Li, Junyou
Ye, Deheng
Lin, Guosheng
Computer Vision and Pattern Recognition
Video-based world models have recently garnered increasing attention for their ability to synthesize diverse and dynamic visual environments. In this paper, we focus on shared world modeling, where a model generates multiple videos from a set of input images, each representing the same underlying world in different camera poses. We propose IC-World, a novel generation framework, enabling parallel generation for all input images via activating the inherent in-context generation capability of large video models. We further finetune IC-World via reinforcement learning, Group Relative Policy Optimization, together with two proposed novel reward models to enforce scene-level geometry consistency and object-level motion consistency among the set of generated videos. Extensive experiments demonstrate that IC-World substantially outperforms state-of-the-art methods in both geometry and motion consistency. To the best of our knowledge, this is the first work to systematically explore the shared world modeling problem with video-based world models.
title IC-World: In-Context Generation for Shared World Modeling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.02793