VideoWorld 2: Learning Transferable Knowledge from Real-world Videos

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ren, Zhongwei, Wei, Yunchao, Yu, Xiao, Luo, Guixun, Zhao, Yao, Kang, Bingyi, Feng, Jiashi, Jin, Xiaojie
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914320199188480
author Ren, Zhongwei
Wei, Yunchao
Yu, Xiao
Luo, Guixun
Zhao, Yao
Kang, Bingyi
Feng, Jiashi
Jin, Xiaojie
author_facet Ren, Zhongwei
Wei, Yunchao
Yu, Xiao
Luo, Guixun
Zhao, Yao
Kang, Bingyi
Feng, Jiashi
Jin, Xiaojie
contents Learning transferable knowledge from unlabeled video data and applying it in new environments is a fundamental capability of intelligent agents. This work presents VideoWorld 2, which extends VideoWorld and offers the first investigation into learning transferable knowledge directly from raw real-world videos. At its core, VideoWorld 2 introduces a dynamic-enhanced Latent Dynamics Model (dLDM) that decouples action dynamics from visual appearance: a pretrained video diffusion model handles visual appearance modeling, enabling the dLDM to learn latent codes that focus on compact and meaningful task-related dynamics. These latent codes are then modeled autoregressively to learn task policies and support long-horizon reasoning. We evaluate VideoWorld 2 on challenging real-world handcraft making tasks, where prior video generation and latent-dynamics models struggle to operate reliably. Remarkably, VideoWorld 2 achieves up to 70% improvement in task success rate and produces coherent long execution videos. In robotics, we show that VideoWorld 2 can acquire effective manipulation knowledge from the Open-X dataset, which substantially improves task performance on CALVIN. This study reveals the potential of learning transferable world knowledge directly from raw videos, with all code, data, and models to be open-sourced for further research.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10102
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
Ren, Zhongwei
Wei, Yunchao
Yu, Xiao
Luo, Guixun
Zhao, Yao
Kang, Bingyi
Feng, Jiashi
Jin, Xiaojie
Computer Vision and Pattern Recognition
Learning transferable knowledge from unlabeled video data and applying it in new environments is a fundamental capability of intelligent agents. This work presents VideoWorld 2, which extends VideoWorld and offers the first investigation into learning transferable knowledge directly from raw real-world videos. At its core, VideoWorld 2 introduces a dynamic-enhanced Latent Dynamics Model (dLDM) that decouples action dynamics from visual appearance: a pretrained video diffusion model handles visual appearance modeling, enabling the dLDM to learn latent codes that focus on compact and meaningful task-related dynamics. These latent codes are then modeled autoregressively to learn task policies and support long-horizon reasoning. We evaluate VideoWorld 2 on challenging real-world handcraft making tasks, where prior video generation and latent-dynamics models struggle to operate reliably. Remarkably, VideoWorld 2 achieves up to 70% improvement in task success rate and produces coherent long execution videos. In robotics, we show that VideoWorld 2 can acquire effective manipulation knowledge from the Open-X dataset, which substantially improves task performance on CALVIN. This study reveals the potential of learning transferable world knowledge directly from raw videos, with all code, data, and models to be open-sourced for further research.
title VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.10102