Chain of World: World Model Thinking in Latent Motion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Fuxiang, Di, Donglin, Tang, Lulu, Zhang, Xuancheng, Fan, Lei, Li, Hao, Wei, Chen, Su, Tonghua, Ma, Baorui
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917310485233664
author Yang, Fuxiang
Di, Donglin
Tang, Lulu
Zhang, Xuancheng
Fan, Lei
Li, Hao
Wei, Chen
Su, Tonghua
Ma, Baorui
author_facet Yang, Fuxiang
Di, Donglin
Tang, Lulu
Zhang, Xuancheng
Fan, Lei
Li, Hao
Wei, Chen
Su, Tonghua
Ma, Baorui
contents Vision-Language-Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynamics. World-model VLAs address this by predicting future frames, but waste capacity reconstructing redundant backgrounds. Latent-action VLAs encode frame-to-frame transitions compactly, but lack temporally continuous dynamic modeling and world knowledge. To overcome these limitations, we introduce CoWVLA (Chain-of-World VLA), a new "Chain of World" paradigm that unifies world-model temporal reasoning with a disentangled latent motion representation. First, a pretrained video VAE serves as a latent motion extractor, explicitly factorizing video segments into structure and motion latents. Then, during pre-training, the VLA learns from an instruction and an initial frame to infer a continuous latent motion chain and predict the segment's terminal frame. Finally, during co-fine-tuning, this latent dynamic is aligned with discrete action prediction by jointly modeling sparse keyframes and action sequences in a unified autoregressive decoder. This design preserves the world-model benefits of temporal reasoning and world knowledge while retaining the compactness and interpretability of latent actions, enabling efficient visuomotor learning. Extensive experiments on robotic simulation benchmarks show that CoWVLA outperforms existing world-model and latent-action approaches and achieves moderate computational efficiency, highlighting its potential as a more effective VLA pretraining paradigm. The project website can be found at https://fx-hit.github.io/cowvla-io.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03195
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Chain of World: World Model Thinking in Latent Motion
Yang, Fuxiang
Di, Donglin
Tang, Lulu
Zhang, Xuancheng
Fan, Lei
Li, Hao
Wei, Chen
Su, Tonghua
Ma, Baorui
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Vision-Language-Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynamics. World-model VLAs address this by predicting future frames, but waste capacity reconstructing redundant backgrounds. Latent-action VLAs encode frame-to-frame transitions compactly, but lack temporally continuous dynamic modeling and world knowledge. To overcome these limitations, we introduce CoWVLA (Chain-of-World VLA), a new "Chain of World" paradigm that unifies world-model temporal reasoning with a disentangled latent motion representation. First, a pretrained video VAE serves as a latent motion extractor, explicitly factorizing video segments into structure and motion latents. Then, during pre-training, the VLA learns from an instruction and an initial frame to infer a continuous latent motion chain and predict the segment's terminal frame. Finally, during co-fine-tuning, this latent dynamic is aligned with discrete action prediction by jointly modeling sparse keyframes and action sequences in a unified autoregressive decoder. This design preserves the world-model benefits of temporal reasoning and world knowledge while retaining the compactness and interpretability of latent actions, enabling efficient visuomotor learning. Extensive experiments on robotic simulation benchmarks show that CoWVLA outperforms existing world-model and latent-action approaches and achieves moderate computational efficiency, highlighting its potential as a more effective VLA pretraining paradigm. The project website can be found at https://fx-hit.github.io/cowvla-io.
title Chain of World: World Model Thinking in Latent Motion
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2603.03195