Learning Emergent Gaits with Decentralized Phase Oscillators: on the role of Observations, Rewards, and Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911778714157056 |
|---|---|
| author | Zhang, Jenny Heim, Steve Jeon, Se Hwan Kim, Sangbae |
| author_facet | Zhang, Jenny Heim, Steve Jeon, Se Hwan Kim, Sangbae |
| contents | We present a minimal phase oscillator model for learning quadrupedal locomotion. Each of the four oscillators is coupled only to itself and its corresponding leg through local feedback of the ground reaction force, which can be interpreted as an observer feedback gain. We interpret the oscillator itself as a latent contact state-estimator. Through a systematic ablation study, we show that the combination of phase observations, simple phase-based rewards, and the local feedback dynamics induces policies that exhibit emergent gait preferences, while using a reduced set of simple rewards, and without prescribing a specific gait. The code is open-source, and a video synopsis available at https://youtu.be/1NKQ0rSV3jU. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_08662 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Learning Emergent Gaits with Decentralized Phase Oscillators: on the role of Observations, Rewards, and Feedback Zhang, Jenny Heim, Steve Jeon, Se Hwan Kim, Sangbae Robotics Machine Learning We present a minimal phase oscillator model for learning quadrupedal locomotion. Each of the four oscillators is coupled only to itself and its corresponding leg through local feedback of the ground reaction force, which can be interpreted as an observer feedback gain. We interpret the oscillator itself as a latent contact state-estimator. Through a systematic ablation study, we show that the combination of phase observations, simple phase-based rewards, and the local feedback dynamics induces policies that exhibit emergent gait preferences, while using a reduced set of simple rewards, and without prescribing a specific gait. The code is open-source, and a video synopsis available at https://youtu.be/1NKQ0rSV3jU. |
| title | Learning Emergent Gaits with Decentralized Phase Oscillators: on the role of Observations, Rewards, and Feedback |
| topic | Robotics Machine Learning |
| url | https://arxiv.org/abs/2402.08662 |