Learning Emergent Gaits with Decentralized Phase Oscillators: on the role of Observations, Rewards, and Feedback

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Jenny, Heim, Steve, Jeon, Se Hwan, Kim, Sangbae
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911778714157056
author Zhang, Jenny
Heim, Steve
Jeon, Se Hwan
Kim, Sangbae
author_facet Zhang, Jenny
Heim, Steve
Jeon, Se Hwan
Kim, Sangbae
contents We present a minimal phase oscillator model for learning quadrupedal locomotion. Each of the four oscillators is coupled only to itself and its corresponding leg through local feedback of the ground reaction force, which can be interpreted as an observer feedback gain. We interpret the oscillator itself as a latent contact state-estimator. Through a systematic ablation study, we show that the combination of phase observations, simple phase-based rewards, and the local feedback dynamics induces policies that exhibit emergent gait preferences, while using a reduced set of simple rewards, and without prescribing a specific gait. The code is open-source, and a video synopsis available at https://youtu.be/1NKQ0rSV3jU.
format Preprint
id arxiv_https___arxiv_org_abs_2402_08662
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Emergent Gaits with Decentralized Phase Oscillators: on the role of Observations, Rewards, and Feedback
Zhang, Jenny
Heim, Steve
Jeon, Se Hwan
Kim, Sangbae
Robotics
Machine Learning
We present a minimal phase oscillator model for learning quadrupedal locomotion. Each of the four oscillators is coupled only to itself and its corresponding leg through local feedback of the ground reaction force, which can be interpreted as an observer feedback gain. We interpret the oscillator itself as a latent contact state-estimator. Through a systematic ablation study, we show that the combination of phase observations, simple phase-based rewards, and the local feedback dynamics induces policies that exhibit emergent gait preferences, while using a reduced set of simple rewards, and without prescribing a specific gait. The code is open-source, and a video synopsis available at https://youtu.be/1NKQ0rSV3jU.
title Learning Emergent Gaits with Decentralized Phase Oscillators: on the role of Observations, Rewards, and Feedback
topic Robotics
Machine Learning
url https://arxiv.org/abs/2402.08662