S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yan, Haodong, Zhong, Zhide, Zhu, Jiaguan, He, Junjie, Yuan, Weilin, Song, Wenxuan, Gong, Xin, Cai, Yingjie, Zhao, Guanyi, Yan, Xu, Liu, Bingbing, Chen, Ying-Cong, Li, Haoang
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910057156837376
author Yan, Haodong
Zhong, Zhide
Zhu, Jiaguan
He, Junjie
Yuan, Weilin
Song, Wenxuan
Gong, Xin
Cai, Yingjie
Zhao, Guanyi
Yan, Xu
Liu, Bingbing
Chen, Ying-Cong
Li, Haoang
author_facet Yan, Haodong
Zhong, Zhide
Zhu, Jiaguan
He, Junjie
Yuan, Weilin
Song, Wenxuan
Gong, Xin
Cai, Yingjie
Zhao, Guanyi
Yan, Xu
Liu, Bingbing
Chen, Ying-Cong
Li, Haoang
contents Video action models (VAMs) have emerged as a promising paradigm for robot learning, owing to their powerful visual foresight for complex manipulation tasks. However, current VAMs, typically relying on either slow multi-step video generation or noisy one-step feature extraction, cannot simultaneously guarantee real-time inference and high-fidelity foresight. To address this limitation, we propose S-VAM, a shortcut video-action model that foresees coherent geometric and semantic representations via a single forward pass. Serving as a stable blueprint, these foreseen representations significantly simplify the action prediction. To enable this efficient shortcut, we introduce a novel self-distillation strategy that condenses structured generative priors of multi-step denoising into one-step inference. Specifically, vision foundation model (VFM) representations extracted from the diffusion model's own multi-step generated videos provide teacher targets. Lightweight decouplers, as students, learn to directly map noisy one-step features to these targets. Extensive experiments in simulation and the real world demonstrate that our S-VAM outperforms state-of-the-art methods, enabling efficient and precise manipulation in complex environments. Our project page is https://haodong-yan.github.io/S-VAM/
format Preprint
id arxiv_https___arxiv_org_abs_2603_16195
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
Yan, Haodong
Zhong, Zhide
Zhu, Jiaguan
He, Junjie
Yuan, Weilin
Song, Wenxuan
Gong, Xin
Cai, Yingjie
Zhao, Guanyi
Yan, Xu
Liu, Bingbing
Chen, Ying-Cong
Li, Haoang
Computer Vision and Pattern Recognition
Robotics
Video action models (VAMs) have emerged as a promising paradigm for robot learning, owing to their powerful visual foresight for complex manipulation tasks. However, current VAMs, typically relying on either slow multi-step video generation or noisy one-step feature extraction, cannot simultaneously guarantee real-time inference and high-fidelity foresight. To address this limitation, we propose S-VAM, a shortcut video-action model that foresees coherent geometric and semantic representations via a single forward pass. Serving as a stable blueprint, these foreseen representations significantly simplify the action prediction. To enable this efficient shortcut, we introduce a novel self-distillation strategy that condenses structured generative priors of multi-step denoising into one-step inference. Specifically, vision foundation model (VFM) representations extracted from the diffusion model's own multi-step generated videos provide teacher targets. Lightweight decouplers, as students, learn to directly map noisy one-step features to these targets. Extensive experiments in simulation and the real world demonstrate that our S-VAM outperforms state-of-the-art methods, enabling efficient and precise manipulation in complex environments. Our project page is https://haodong-yan.github.io/S-VAM/
title S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2603.16195