Saved in:
Bibliographic Details
Main Authors: Sun, Yirui, Zhuge, Guangyu, Liu, Keliang, Gu, Jie, Bing, Xinyu, Gan, Zhongxue, Tian, Chunxu
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.27947
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910264615501824
author Sun, Yirui
Zhuge, Guangyu
Liu, Keliang
Gu, Jie
Bing, Xinyu
Gan, Zhongxue
Tian, Chunxu
author_facet Sun, Yirui
Zhuge, Guangyu
Liu, Keliang
Gu, Jie
Bing, Xinyu
Gan, Zhongxue
Tian, Chunxu
contents World Action Models (WAMs) improve robot manipulation by using video-based future representations to condition action generation. In pixel-space WAMs, however, the best action condition is not necessarily the fully denoised video. Controlled denoising-depth scans show that video refinement can reduce action error up to a state-dependent point, after which the gain may saturate or even reverse when late predictions become less action-relevant or physically unreliable. This suggests that action generation should use a state-dependent point along the video noise trajectory rather than a fixed terminal denoising depth. We introduce State-Adaptive Noise Trajectory Scheduler (SANTS), a lightweight scheduler for video-to-action diffusion policies. At each video decision point, SANTS reads the current video-state representation and noise level, then jointly predicts a cumulative stopping hazard and a relative noise-progression ratio. SANTS is post-trained with a path-level reward computed after the frozen action branch generates the final action chunk, so the scheduler is optimized for downstream action quality rather than intermediate video fidelity, while redundant video-state updates are explicitly penalized. Experiments show that SANTS reaches \(94.4\%\) overall success on RoboTwin 2.0 and \(73.1\%\) average success across seven real-robot tasks, while reducing latency by \(81.7\%\) and \(79.0\%\) relative to full video denoising, respectively. These results indicate that adaptive selection along the video noise trajectory can preserve the control benefits of WAM-style future reasoning while removing much of its redundant inference cost.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27947
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SANTS: A State-Adaptive Scheduler for World Action Models
Sun, Yirui
Zhuge, Guangyu
Liu, Keliang
Gu, Jie
Bing, Xinyu
Gan, Zhongxue
Tian, Chunxu
Robotics
World Action Models (WAMs) improve robot manipulation by using video-based future representations to condition action generation. In pixel-space WAMs, however, the best action condition is not necessarily the fully denoised video. Controlled denoising-depth scans show that video refinement can reduce action error up to a state-dependent point, after which the gain may saturate or even reverse when late predictions become less action-relevant or physically unreliable. This suggests that action generation should use a state-dependent point along the video noise trajectory rather than a fixed terminal denoising depth. We introduce State-Adaptive Noise Trajectory Scheduler (SANTS), a lightweight scheduler for video-to-action diffusion policies. At each video decision point, SANTS reads the current video-state representation and noise level, then jointly predicts a cumulative stopping hazard and a relative noise-progression ratio. SANTS is post-trained with a path-level reward computed after the frozen action branch generates the final action chunk, so the scheduler is optimized for downstream action quality rather than intermediate video fidelity, while redundant video-state updates are explicitly penalized. Experiments show that SANTS reaches \(94.4\%\) overall success on RoboTwin 2.0 and \(73.1\%\) average success across seven real-robot tasks, while reducing latency by \(81.7\%\) and \(79.0\%\) relative to full video denoising, respectively. These results indicate that adaptive selection along the video noise trajectory can preserve the control benefits of WAM-style future reasoning while removing much of its redundant inference cost.
title SANTS: A State-Adaptive Scheduler for World Action Models
topic Robotics
url https://arxiv.org/abs/2605.27947