Actuator-Constrained Reinforcement Learning for High-Speed Quadrupedal Locomotion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shin, Young-Ha, Song, Tae-Gyu, Ji, Gwanghyeon, Park, Hae-Won
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910283533910016
author Shin, Young-Ha
Song, Tae-Gyu
Ji, Gwanghyeon
Park, Hae-Won
author_facet Shin, Young-Ha
Song, Tae-Gyu
Ji, Gwanghyeon
Park, Hae-Won
contents This paper presents a method for achieving high-speed running of a quadruped robot by considering the actuator torque-speed operating region in reinforcement learning. The physical properties and constraints of the actuator are included in the training process to reduce state transitions that are infeasible in the real world due to motor torque-speed limitations. The gait reward is designed to distribute motor torque evenly across all legs, contributing to more balanced power usage and mitigating performance bottlenecks due to single-motor saturation. Additionally, we designed a lightweight foot to enhance the robot's agility. We observed that applying the motor operating region as a constraint helps the policy network avoid infeasible areas during sampling. With the trained policy, KAIST Hound, a 45 kg quadruped robot, can run up to 6.5 m/s, which is the fastest speed among electric motor-based quadruped robots.
format Preprint
id arxiv_https___arxiv_org_abs_2312_17507
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Actuator-Constrained Reinforcement Learning for High-Speed Quadrupedal Locomotion
Shin, Young-Ha
Song, Tae-Gyu
Ji, Gwanghyeon
Park, Hae-Won
Robotics
This paper presents a method for achieving high-speed running of a quadruped robot by considering the actuator torque-speed operating region in reinforcement learning. The physical properties and constraints of the actuator are included in the training process to reduce state transitions that are infeasible in the real world due to motor torque-speed limitations. The gait reward is designed to distribute motor torque evenly across all legs, contributing to more balanced power usage and mitigating performance bottlenecks due to single-motor saturation. Additionally, we designed a lightweight foot to enhance the robot's agility. We observed that applying the motor operating region as a constraint helps the policy network avoid infeasible areas during sampling. With the trained policy, KAIST Hound, a 45 kg quadruped robot, can run up to 6.5 m/s, which is the fastest speed among electric motor-based quadruped robots.
title Actuator-Constrained Reinforcement Learning for High-Speed Quadrupedal Locomotion
topic Robotics
url https://arxiv.org/abs/2312.17507