OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yushan, Sun, Peibo, Li, Shoujie, Xie, Yifan, Zhang, Lingfeng, Chao, Xintao, Dong, Shiyuan, Chen, Fang, Zhang, Xiao-Ping, Ding, Wenbo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914539817140224
author Liu, Yushan
Sun, Peibo
Li, Shoujie
Xie, Yifan
Zhang, Lingfeng
Chao, Xintao
Dong, Shiyuan
Chen, Fang
Zhang, Xiao-Ping
Ding, Wenbo
author_facet Liu, Yushan
Sun, Peibo
Li, Shoujie
Xie, Yifan
Zhang, Lingfeng
Chao, Xintao
Dong, Shiyuan
Chen, Fang
Zhang, Xiao-Ping
Ding, Wenbo
contents World Action Models (WAMs) enhance Vision-Language-Action policies by jointly predicting scene evolution and robot actions, but existing methods usually represent the predicted world as holistic images, video tokens, or global latents. These representations are difficult for an action decoder to address when an instruction refers to a particular object, especially under scene shifts where object identity is entangled with context. We propose OA-WAM, an Object-Addressable World Action Model for robust robot manipulation. OA-WAM decomposes each frame into N+1 slot states, with one robot slot and N object slots. Each slot contains a persistent address vector and a time-varying content vector, and is fused with text, image, proprioception, and past-action tokens in a block-causal sequence. A world head predicts next-frame slot states, while a flow-matching action head decodes a 16-step continuous action chunk in the same forward pass. Addressability is enforced by routing cross-slot attention through address-only keys and resetting the address slice at every transformer layer, separating which object to act on from what that object currently is without adding extra tokens. OA-WAM matches strong VLA and WAM baselines on LIBERO (97.8%) and SimplerEnv (79.3%), reaches state-of-the-art performance on the most relevant LIBERO-Plus geometric axes, and remains competitive on the seven-axis aggregate. A causal slot-intervention test yields a swap-binding cosine of 0.87, versus at most 0.09 for holistic baselines. These results suggest that addressable object states provide an effective interface for robust world-action modeling under scene perturbations.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06481
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation
Liu, Yushan
Sun, Peibo
Li, Shoujie
Xie, Yifan
Zhang, Lingfeng
Chao, Xintao
Dong, Shiyuan
Chen, Fang
Zhang, Xiao-Ping
Ding, Wenbo
Robotics
World Action Models (WAMs) enhance Vision-Language-Action policies by jointly predicting scene evolution and robot actions, but existing methods usually represent the predicted world as holistic images, video tokens, or global latents. These representations are difficult for an action decoder to address when an instruction refers to a particular object, especially under scene shifts where object identity is entangled with context. We propose OA-WAM, an Object-Addressable World Action Model for robust robot manipulation. OA-WAM decomposes each frame into N+1 slot states, with one robot slot and N object slots. Each slot contains a persistent address vector and a time-varying content vector, and is fused with text, image, proprioception, and past-action tokens in a block-causal sequence. A world head predicts next-frame slot states, while a flow-matching action head decodes a 16-step continuous action chunk in the same forward pass. Addressability is enforced by routing cross-slot attention through address-only keys and resetting the address slice at every transformer layer, separating which object to act on from what that object currently is without adding extra tokens. OA-WAM matches strong VLA and WAM baselines on LIBERO (97.8%) and SimplerEnv (79.3%), reaches state-of-the-art performance on the most relevant LIBERO-Plus geometric axes, and remains competitive on the seven-axis aggregate. A causal slot-intervention test yields a swap-binding cosine of 0.87, versus at most 0.09 for holistic baselines. These results suggest that addressable object states provide an effective interface for robust world-action modeling under scene perturbations.
title OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation
topic Robotics
url https://arxiv.org/abs/2605.06481