Learning Primitive Embodied World Models: Towards Scalable Robotic Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Qiao, Yang, Liujia, Tang, Wei, Huang, Wei, Xu, Kaixin, Chen, Yongchao, Liu, Mingyu, Yang, Jiange, Zhu, Haoyi, Wang, Yating, He, Tong, Chen, Yilun, Dai, Xili, Ye, Nanyang, Gu, Qinying
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918215772274688
author Sun, Qiao
Yang, Liujia
Tang, Wei
Huang, Wei
Xu, Kaixin
Chen, Yongchao
Liu, Mingyu
Yang, Jiange
Zhu, Haoyi
Wang, Yating
He, Tong
Chen, Yilun
Dai, Xili
Ye, Nanyang
Gu, Qinying
author_facet Sun, Qiao
Yang, Liujia
Tang, Wei
Huang, Wei
Xu, Kaixin
Chen, Yongchao
Liu, Mingyu
Yang, Jiange
Zhu, Haoyi
Wang, Yating
He, Tong
Chen, Yilun
Dai, Xili
Ye, Nanyang
Gu, Qinying
contents While video-generation-based embodied world models have gained increasing attention, their reliance on large-scale embodied interaction data remains a key bottleneck. The scarcity, difficulty of collection, and high dimensionality of embodied data fundamentally limit the alignment granularity between language and actions and exacerbate the challenge of long-horizon video generation--hindering generative models from achieving a "GPT moment" in the embodied domain. There is a naive observation: the diversity of embodied data far exceeds the relatively small space of possible primitive motions. Based on this insight, we propose a novel paradigm for world modeling--Primitive Embodied World Models (PEWM). By restricting video generation to fixed short horizons, our approach 1) enables fine-grained alignment between linguistic concepts and visual representations of robotic actions, 2) reduces learning complexity, 3) improves data efficiency in embodied data collection, and 4) decreases inference latency. By equipping with a modular Vision-Language Model (VLM) planner and a Start-Goal heatmap Guidance mechanism (SGG), PEWM further enables flexible closed-loop control and supports compositional generalization of primitive-level policies over extended, complex tasks. Our framework leverages the spatiotemporal vision priors in video models and the semantic awareness of VLMs to bridge the gap between fine-grained physical interaction and high-level reasoning, paving the way toward scalable, interpretable, and general-purpose embodied intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20840
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
Sun, Qiao
Yang, Liujia
Tang, Wei
Huang, Wei
Xu, Kaixin
Chen, Yongchao
Liu, Mingyu
Yang, Jiange
Zhu, Haoyi
Wang, Yating
He, Tong
Chen, Yilun
Dai, Xili
Ye, Nanyang
Gu, Qinying
Robotics
Artificial Intelligence
Multimedia
While video-generation-based embodied world models have gained increasing attention, their reliance on large-scale embodied interaction data remains a key bottleneck. The scarcity, difficulty of collection, and high dimensionality of embodied data fundamentally limit the alignment granularity between language and actions and exacerbate the challenge of long-horizon video generation--hindering generative models from achieving a "GPT moment" in the embodied domain. There is a naive observation: the diversity of embodied data far exceeds the relatively small space of possible primitive motions. Based on this insight, we propose a novel paradigm for world modeling--Primitive Embodied World Models (PEWM). By restricting video generation to fixed short horizons, our approach 1) enables fine-grained alignment between linguistic concepts and visual representations of robotic actions, 2) reduces learning complexity, 3) improves data efficiency in embodied data collection, and 4) decreases inference latency. By equipping with a modular Vision-Language Model (VLM) planner and a Start-Goal heatmap Guidance mechanism (SGG), PEWM further enables flexible closed-loop control and supports compositional generalization of primitive-level policies over extended, complex tasks. Our framework leverages the spatiotemporal vision priors in video models and the semantic awareness of VLMs to bridge the gap between fine-grained physical interaction and high-level reasoning, paving the way toward scalable, interpretable, and general-purpose embodied intelligence.
title Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
topic Robotics
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2508.20840