StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Mingyu, Shu, Jiuhe, Chen, Hui, Li, Zeju, Zhao, Canyu, Yang, Jiange, Gao, Shenyuan, Chen, Hao, Shen, Chunhua
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918440198995968
author Liu, Mingyu
Shu, Jiuhe
Chen, Hui
Li, Zeju
Zhao, Canyu
Yang, Jiange
Gao, Shenyuan
Chen, Hao
Shen, Chunhua
author_facet Liu, Mingyu
Shu, Jiuhe
Chen, Hui
Li, Zeju
Zhao, Canyu
Yang, Jiange
Gao, Shenyuan
Chen, Hao
Shen, Chunhua
contents A fundamental challenge in embodied intelligence is developing expressive and compact state representations for efficient world modeling and decision making. However, existing methods often fail to achieve this balance, yielding representations that are either overly redundant or lacking in task-critical information. We propose an unsupervised approach that learns a highly compressed two-token state representation using a lightweight encoder and a pre-trained Diffusion Transformer (DiT) decoder, capitalizing on its strong generative prior. Our representation is efficient, interpretable, and integrates seamlessly into existing VLA-based models, improving performance by 14.3% on LIBERO and 30% in real-world task success with minimal inference overhead. More importantly, we find that the difference between these tokens, obtained via latent interpolation, naturally serves as a highly effective latent action, which can be further decoded into executable robot actions. This emergent capability reveals that our representation captures structured dynamics without explicit supervision. We name our method StaMo for its ability to learn generalizable robotic Motion from compact State representation, which is encoded from static images, challenging the prevalent dependence to learning latent action on complex architectures and video data. The resulting latent actions also enhance policy co-training, outperforming prior methods by 10.4% with improved interpretability. Moreover, our approach scales effectively across diverse data sources, including real-world robot data, simulation, and human egocentric video.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05057
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
Liu, Mingyu
Shu, Jiuhe
Chen, Hui
Li, Zeju
Zhao, Canyu
Yang, Jiange
Gao, Shenyuan
Chen, Hao
Shen, Chunhua
Robotics
Computer Vision and Pattern Recognition
A fundamental challenge in embodied intelligence is developing expressive and compact state representations for efficient world modeling and decision making. However, existing methods often fail to achieve this balance, yielding representations that are either overly redundant or lacking in task-critical information. We propose an unsupervised approach that learns a highly compressed two-token state representation using a lightweight encoder and a pre-trained Diffusion Transformer (DiT) decoder, capitalizing on its strong generative prior. Our representation is efficient, interpretable, and integrates seamlessly into existing VLA-based models, improving performance by 14.3% on LIBERO and 30% in real-world task success with minimal inference overhead. More importantly, we find that the difference between these tokens, obtained via latent interpolation, naturally serves as a highly effective latent action, which can be further decoded into executable robot actions. This emergent capability reveals that our representation captures structured dynamics without explicit supervision. We name our method StaMo for its ability to learn generalizable robotic Motion from compact State representation, which is encoded from static images, challenging the prevalent dependence to learning latent action on complex architectures and video data. The resulting latent actions also enhance policy co-training, outperforming prior methods by 10.4% with improved interpretability. Moreover, our approach scales effectively across diverse data sources, including real-world robot data, simulation, and human egocentric video.
title StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.05057