From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhe, Chi, Cheng, Wei, Yangyang, Zhu, Boan, Peng, Yibo, Huang, Tao, Wang, Pengwei, Wang, Zhongyuan, Zhang, Shanghang, Xu, Chang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912654964031488
author Li, Zhe
Chi, Cheng
Wei, Yangyang
Zhu, Boan
Peng, Yibo
Huang, Tao
Wang, Pengwei
Wang, Zhongyuan
Zhang, Shanghang
Xu, Chang
author_facet Li, Zhe
Chi, Cheng
Wei, Yangyang
Zhu, Boan
Peng, Yibo
Huang, Tao
Wang, Pengwei
Wang, Zhongyuan
Zhang, Shanghang
Xu, Chang
contents Natural language offers a natural interface for humanoid robots, but existing language-guided humanoid locomotion pipelines remain cumbersome and untrustworthy. They typically decode human motion, retarget it to robot morphology, and then track it with a physics-based controller. However, this multi-stage process is prone to cumulative errors, introduces high latency, and yields weak coupling between semantics and control. These limitations call for a more direct pathway from language to action, one that eliminates fragile intermediate stages. Therefore, we present RoboGhost, a retargeting-free framework that directly conditions humanoid policies on language-grounded motion latents. By bypassing explicit motion decoding and retargeting, RoboGhost enables a diffusion-based policy to denoise executable actions directly from noise, preserving semantic intent and supporting fast, reactive control. A hybrid causal transformer-diffusion motion generator further ensures long-horizon consistency while maintaining stability and diversity, yielding rich latent representations for precise humanoid behavior. Extensive experiments demonstrate that RoboGhost substantially reduces deployment latency, improves success rates and tracking precision, and produces smooth, semantically aligned locomotion on real humanoids. Beyond text, the framework naturally extends to other modalities such as images, audio, and music, providing a universal foundation for vision-language-action humanoid systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14952
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance
Li, Zhe
Chi, Cheng
Wei, Yangyang
Zhu, Boan
Peng, Yibo
Huang, Tao
Wang, Pengwei
Wang, Zhongyuan
Zhang, Shanghang
Xu, Chang
Robotics
Computer Vision and Pattern Recognition
Natural language offers a natural interface for humanoid robots, but existing language-guided humanoid locomotion pipelines remain cumbersome and untrustworthy. They typically decode human motion, retarget it to robot morphology, and then track it with a physics-based controller. However, this multi-stage process is prone to cumulative errors, introduces high latency, and yields weak coupling between semantics and control. These limitations call for a more direct pathway from language to action, one that eliminates fragile intermediate stages. Therefore, we present RoboGhost, a retargeting-free framework that directly conditions humanoid policies on language-grounded motion latents. By bypassing explicit motion decoding and retargeting, RoboGhost enables a diffusion-based policy to denoise executable actions directly from noise, preserving semantic intent and supporting fast, reactive control. A hybrid causal transformer-diffusion motion generator further ensures long-horizon consistency while maintaining stability and diversity, yielding rich latent representations for precise humanoid behavior. Extensive experiments demonstrate that RoboGhost substantially reduces deployment latency, improves success rates and tracking precision, and produces smooth, semantically aligned locomotion on real humanoids. Beyond text, the framework naturally extends to other modalities such as images, audio, and music, providing a universal foundation for vision-language-action humanoid systems.
title From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.14952