Lifting Embodied World Models for Planning and Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Alex N., Darrell, Trevor, Izmailov, Pavel, Bai, Yutong, Bar, Amir
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918472907227136
author Wang, Alex N.
Darrell, Trevor
Izmailov, Pavel
Bai, Yutong
Bar, Amir
author_facet Wang, Alex N.
Darrell, Trevor
Izmailov, Pavel
Bai, Yutong
Bar, Amir
contents World models of embodied agents predict future observations conditioned on an action taken by the agent. For complex embodiments, action spaces are high-dimensional and difficult to specify: for example, precisely controlling a human agent requires specifying the motion of each joint. This makes the world model hard to control and expensive to plan with as search-based methods like CEM scale poorly with action dimensionality. To address this issue, we train a lightweight policy that maps high-level actions to sequences of low-level joint actions. Composing this policy with the frozen world model produces a lifted world model that predicts a sequence of future observations from a single high-level action. We instantiate this framework for a human-like embodiment, defining the high-level action space as a small set of 2D waypoints annotated on the current observation frame, each specifying a near-term goal position for a leaf joint (pelvis, head, hands). Waypoints are low-dimensional, visually interpretable, and easy to specify manually or to search over. We show that the lifted world model substantially outperforms searching directly in low-level joint space ($3.8\times$ lower mean joint error to the goal pose), while remaining more compute-efficient and generalizing to environments unseen by the policy.
format Preprint
id arxiv_https___arxiv_org_abs_2604_26182
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Lifting Embodied World Models for Planning and Control
Wang, Alex N.
Darrell, Trevor
Izmailov, Pavel
Bai, Yutong
Bar, Amir
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
World models of embodied agents predict future observations conditioned on an action taken by the agent. For complex embodiments, action spaces are high-dimensional and difficult to specify: for example, precisely controlling a human agent requires specifying the motion of each joint. This makes the world model hard to control and expensive to plan with as search-based methods like CEM scale poorly with action dimensionality. To address this issue, we train a lightweight policy that maps high-level actions to sequences of low-level joint actions. Composing this policy with the frozen world model produces a lifted world model that predicts a sequence of future observations from a single high-level action. We instantiate this framework for a human-like embodiment, defining the high-level action space as a small set of 2D waypoints annotated on the current observation frame, each specifying a near-term goal position for a leaf joint (pelvis, head, hands). Waypoints are low-dimensional, visually interpretable, and easy to specify manually or to search over. We show that the lifted world model substantially outperforms searching directly in low-level joint space ($3.8\times$ lower mean joint error to the goal pose), while remaining more compute-efficient and generalizing to environments unseen by the policy.
title Lifting Embodied World Models for Planning and Control
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2604.26182