Gespeichert in:
| Hauptverfasser: | , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2512.02793 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866908687711338496 |
|---|---|
| author | Wu, Fan Wei, Jiacheng Li, Ruibo Xu, Yi Li, Junyou Ye, Deheng Lin, Guosheng |
| author_facet | Wu, Fan Wei, Jiacheng Li, Ruibo Xu, Yi Li, Junyou Ye, Deheng Lin, Guosheng |
| contents | Video-based world models have recently garnered increasing attention for their ability to synthesize diverse and dynamic visual environments. In this paper, we focus on shared world modeling, where a model generates multiple videos from a set of input images, each representing the same underlying world in different camera poses. We propose IC-World, a novel generation framework, enabling parallel generation for all input images via activating the inherent in-context generation capability of large video models. We further finetune IC-World via reinforcement learning, Group Relative Policy Optimization, together with two proposed novel reward models to enforce scene-level geometry consistency and object-level motion consistency among the set of generated videos. Extensive experiments demonstrate that IC-World substantially outperforms state-of-the-art methods in both geometry and motion consistency. To the best of our knowledge, this is the first work to systematically explore the shared world modeling problem with video-based world models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_02793 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | IC-World: In-Context Generation for Shared World Modeling Wu, Fan Wei, Jiacheng Li, Ruibo Xu, Yi Li, Junyou Ye, Deheng Lin, Guosheng Computer Vision and Pattern Recognition Video-based world models have recently garnered increasing attention for their ability to synthesize diverse and dynamic visual environments. In this paper, we focus on shared world modeling, where a model generates multiple videos from a set of input images, each representing the same underlying world in different camera poses. We propose IC-World, a novel generation framework, enabling parallel generation for all input images via activating the inherent in-context generation capability of large video models. We further finetune IC-World via reinforcement learning, Group Relative Policy Optimization, together with two proposed novel reward models to enforce scene-level geometry consistency and object-level motion consistency among the set of generated videos. Extensive experiments demonstrate that IC-World substantially outperforms state-of-the-art methods in both geometry and motion consistency. To the best of our knowledge, this is the first work to systematically explore the shared world modeling problem with video-based world models. |
| title | IC-World: In-Context Generation for Shared World Modeling |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.02793 |