When to Trust Imagination: Adaptive Action Execution for World Action Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Rui, Zhang, Yue, Lin, Jiehong, Luo, Kuncheng, Wang, Jianan, Wang, Zhongrui, Qi, Xiaojuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909029387730944
author Wang, Rui
Zhang, Yue
Lin, Jiehong
Luo, Kuncheng
Wang, Jianan
Wang, Zhongrui
Qi, Xiaojuan
author_facet Wang, Rui
Zhang, Yue
Lin, Jiehong
Luo, Kuncheng
Wang, Jianan
Wang, Zhongrui
Qi, Xiaojuan
contents World Action Models (WAMs) have recently emerged as a promising paradigm for robotic manipulation by jointly predicting future visual observations and future actions. However, current WAMs typically execute a fixed number of predicted actions after each model inference, leaving the robot blind to whether the imagined future remains consistent with the actual physical rollout. In this work, we formulate adaptive WAM execution as a future-reality verification problem: the robot should execute longer when the WAM-predicted future remains reliable, and replan earlier when reality deviates from imagination. To this end, we propose Future Forward Dynamics Causal Attention (FFDC), a lightweight verifier that jointly reasons over predicted future actions, predicted visual dynamics, real observations, and language instructions to estimate whether the remaining action rollout can still be trusted. FFDC enables adaptive action chunk sizes as an emergent consequence of prediction-observation consistency, preserving the efficiency of long-horizon execution while restoring responsiveness in contact-rich or difficult phases. We further introduce Mixture-of-Horizon Training to improve long-horizon trajectory coverage for adaptive execution. Experiments on the RoboTwin benchmark and in the real world demonstrate that our method achieves a strong robustness-efficiency trade-off: on RoboTwin, it reduces WAM forward passes by 69.10% and execution time by 34.02%, while improving success rate by 2.54% over the short-chunk baseline; in real-world experiments, it improves success rate by 35%.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06222
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When to Trust Imagination: Adaptive Action Execution for World Action Models
Wang, Rui
Zhang, Yue
Lin, Jiehong
Luo, Kuncheng
Wang, Jianan
Wang, Zhongrui
Qi, Xiaojuan
Robotics
Artificial Intelligence
World Action Models (WAMs) have recently emerged as a promising paradigm for robotic manipulation by jointly predicting future visual observations and future actions. However, current WAMs typically execute a fixed number of predicted actions after each model inference, leaving the robot blind to whether the imagined future remains consistent with the actual physical rollout. In this work, we formulate adaptive WAM execution as a future-reality verification problem: the robot should execute longer when the WAM-predicted future remains reliable, and replan earlier when reality deviates from imagination. To this end, we propose Future Forward Dynamics Causal Attention (FFDC), a lightweight verifier that jointly reasons over predicted future actions, predicted visual dynamics, real observations, and language instructions to estimate whether the remaining action rollout can still be trusted. FFDC enables adaptive action chunk sizes as an emergent consequence of prediction-observation consistency, preserving the efficiency of long-horizon execution while restoring responsiveness in contact-rich or difficult phases. We further introduce Mixture-of-Horizon Training to improve long-horizon trajectory coverage for adaptive execution. Experiments on the RoboTwin benchmark and in the real world demonstrate that our method achieves a strong robustness-efficiency trade-off: on RoboTwin, it reduces WAM forward passes by 69.10% and execution time by 34.02%, while improving success rate by 2.54% over the short-chunk baseline; in real-world experiments, it improves success rate by 35%.
title When to Trust Imagination: Adaptive Action Execution for World Action Models
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2605.06222