Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jin, Minghao, Liao, Mozheng, Han, Mingfei, Li, Zhihui, Chang, Xiaojun
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908883453214720
author Jin, Minghao
Liao, Mozheng
Han, Mingfei
Li, Zhihui
Chang, Xiaojun
author_facet Jin, Minghao
Liao, Mozheng
Han, Mingfei
Li, Zhihui
Chang, Xiaojun
contents Recent world-model-based Vision-Language-Action (VLA) architectures have improved robotic manipulation through predictive visual foresight. However, dense future prediction introduces visual redundancy and accumulates errors, causing long-horizon plan drift. Meanwhile, recent sparse methods typically represent visual foresight using high-level semantic subtasks or implicit latent states. These representations often lack explicit kinematic grounding, weakening the alignment between planning and low-level execution. To address this, we propose StructVLA, which reformulates a generative world model into an explicit structured planner for reliable control. Instead of dense rollouts or semantic goals, StructVLA predicts sparse, physically meaningful structured frames. Derived from intrinsic kinematic cues (e.g., gripper transitions and kinematic turning points), these frames capture spatiotemporal milestones closely aligned with task progress. We implement this approach through a two-stage training paradigm with a unified discrete token vocabulary: the world model is first trained to predict structured frames and subsequently optimized to map the structured foresight into low-level actions. This approach provides clear physical guidance and bridges visual planning and motion control. In our experiments, StructVLA achieves strong average success rates of 75.0% on SimplerEnv-WidowX and 94.8% on LIBERO. Real-world deployments further demonstrate reliable task completion and robust generalization across both basic pick-and-place and complex long-horizon tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12553
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation
Jin, Minghao
Liao, Mozheng
Han, Mingfei
Li, Zhihui
Chang, Xiaojun
Robotics
Computer Vision and Pattern Recognition
Recent world-model-based Vision-Language-Action (VLA) architectures have improved robotic manipulation through predictive visual foresight. However, dense future prediction introduces visual redundancy and accumulates errors, causing long-horizon plan drift. Meanwhile, recent sparse methods typically represent visual foresight using high-level semantic subtasks or implicit latent states. These representations often lack explicit kinematic grounding, weakening the alignment between planning and low-level execution. To address this, we propose StructVLA, which reformulates a generative world model into an explicit structured planner for reliable control. Instead of dense rollouts or semantic goals, StructVLA predicts sparse, physically meaningful structured frames. Derived from intrinsic kinematic cues (e.g., gripper transitions and kinematic turning points), these frames capture spatiotemporal milestones closely aligned with task progress. We implement this approach through a two-stage training paradigm with a unified discrete token vocabulary: the world model is first trained to predict structured frames and subsequently optimized to map the structured foresight into low-level actions. This approach provides clear physical guidance and bridges visual planning and motion control. In our experiments, StructVLA achieves strong average success rates of 75.0% on SimplerEnv-WidowX and 94.8% on LIBERO. Real-world deployments further demonstrate reliable task completion and robust generalization across both basic pick-and-place and complex long-horizon tasks.
title Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.12553