LocoVLM: Grounding Vision and Language for Adapting Versatile Legged Locomotion Policies

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nahrendra, I Made Aswin, Lee, Seunghyun, Lee, Dongkyu, Myung, Hyun
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918331972321280
author Nahrendra, I Made Aswin
Lee, Seunghyun
Lee, Dongkyu
Myung, Hyun
author_facet Nahrendra, I Made Aswin
Lee, Seunghyun
Lee, Dongkyu
Myung, Hyun
contents Recent advances in legged locomotion learning are still dominated by the utilization of geometric representations of the environment, limiting the robot's capability to respond to higher-level semantics such as human instructions. To address this limitation, we propose a novel approach that integrates high-level commonsense reasoning from foundation models into the process of legged locomotion adaptation. Specifically, our method utilizes a pre-trained large language model to synthesize an instruction-grounded skill database tailored for legged robots. A pre-trained vision-language model is employed to extract high-level environmental semantics and ground them within the skill database, enabling real-time skill advisories for the robot. To facilitate versatile skill control, we train a style-conditioned policy capable of generating diverse and robust locomotion skills with high fidelity to specified styles. To the best of our knowledge, this is the first work to demonstrate real-time adaptation of legged locomotion using high-level reasoning from environmental semantics and instructions with instruction-following accuracy of up to 87% without the need for online query to on-the-cloud foundation models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10399
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LocoVLM: Grounding Vision and Language for Adapting Versatile Legged Locomotion Policies
Nahrendra, I Made Aswin
Lee, Seunghyun
Lee, Dongkyu
Myung, Hyun
Robotics
Recent advances in legged locomotion learning are still dominated by the utilization of geometric representations of the environment, limiting the robot's capability to respond to higher-level semantics such as human instructions. To address this limitation, we propose a novel approach that integrates high-level commonsense reasoning from foundation models into the process of legged locomotion adaptation. Specifically, our method utilizes a pre-trained large language model to synthesize an instruction-grounded skill database tailored for legged robots. A pre-trained vision-language model is employed to extract high-level environmental semantics and ground them within the skill database, enabling real-time skill advisories for the robot. To facilitate versatile skill control, we train a style-conditioned policy capable of generating diverse and robust locomotion skills with high fidelity to specified styles. To the best of our knowledge, this is the first work to demonstrate real-time adaptation of legged locomotion using high-level reasoning from environmental semantics and instructions with instruction-following accuracy of up to 87% without the need for online query to on-the-cloud foundation models.
title LocoVLM: Grounding Vision and Language for Adapting Versatile Legged Locomotion Policies
topic Robotics
url https://arxiv.org/abs/2602.10399