High-Performance Reinforcement Learning on Spot: Optimizing Simulation Parameters with Distributional Measures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Miller, AJ, Yu, Fangzhou, Brauckmann, Michael, Farshidian, Farbod
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915369968467968
author Miller, AJ
Yu, Fangzhou
Brauckmann, Michael
Farshidian, Farbod
author_facet Miller, AJ
Yu, Fangzhou
Brauckmann, Michael
Farshidian, Farbod
contents This work presents an overview of the technical details behind a high performance reinforcement learning policy deployment with the Spot RL Researcher Development Kit for low level motor access on Boston Dynamics Spot. This represents the first public demonstration of an end to end end reinforcement learning policy deployed on Spot hardware with training code publicly available through Nvidia IsaacLab and deployment code available through Boston Dynamics. We utilize Wasserstein Distance and Maximum Mean Discrepancy to quantify the distributional dissimilarity of data collected on hardware and in simulation to measure our sim2real gap. We use these measures as a scoring function for the Covariance Matrix Adaptation Evolution Strategy to optimize simulated parameters that are unknown or difficult to measure from Spot. Our procedure for modeling and training produces high quality reinforcement learning policies capable of multiple gaits, including a flight phase. We deploy policies capable of over 5.2ms locomotion, more than triple Spots default controller maximum speed, robustness to slippery surfaces, disturbance rejection, and overall agility previously unseen on Spot. We detail our method and release our code to support future work on Spot with the low level API.
format Preprint
id arxiv_https___arxiv_org_abs_2504_17857
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle High-Performance Reinforcement Learning on Spot: Optimizing Simulation Parameters with Distributional Measures
Miller, AJ
Yu, Fangzhou
Brauckmann, Michael
Farshidian, Farbod
Machine Learning
Robotics
This work presents an overview of the technical details behind a high performance reinforcement learning policy deployment with the Spot RL Researcher Development Kit for low level motor access on Boston Dynamics Spot. This represents the first public demonstration of an end to end end reinforcement learning policy deployed on Spot hardware with training code publicly available through Nvidia IsaacLab and deployment code available through Boston Dynamics. We utilize Wasserstein Distance and Maximum Mean Discrepancy to quantify the distributional dissimilarity of data collected on hardware and in simulation to measure our sim2real gap. We use these measures as a scoring function for the Covariance Matrix Adaptation Evolution Strategy to optimize simulated parameters that are unknown or difficult to measure from Spot. Our procedure for modeling and training produces high quality reinforcement learning policies capable of multiple gaits, including a flight phase. We deploy policies capable of over 5.2ms locomotion, more than triple Spots default controller maximum speed, robustness to slippery surfaces, disturbance rejection, and overall agility previously unseen on Spot. We detail our method and release our code to support future work on Spot with the low level API.
title High-Performance Reinforcement Learning on Spot: Optimizing Simulation Parameters with Distributional Measures
topic Machine Learning
Robotics
url https://arxiv.org/abs/2504.17857