Saved in:
Bibliographic Details
Main Authors: Liu, Yichen, Viswanadha, Kesava, Li, Zhongyu, Lojo, Nelson, Pister, Kristofer S. J.
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.24740
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912797829365760
author Liu, Yichen
Viswanadha, Kesava
Li, Zhongyu
Lojo, Nelson
Pister, Kristofer S. J.
author_facet Liu, Yichen
Viswanadha, Kesava
Li, Zhongyu
Lojo, Nelson
Pister, Kristofer S. J.
contents An important function of autonomous microrobots is the ability to perform robust movement over terrain. This paper explores an edge ML approach to microrobot locomotion, allowing for on-device, lower latency control under compute, memory, and power constraints. This paper explores the locomotion of a sub-centimeter quadrupedal microrobot via reinforcement learning (RL) and deploys the resulting controller on an ultra-small system-on-chip (SoC), SC$μ$M-3C, featuring an ARM Cortex-M0 microcontroller running at 5 MHz. We train a compact FP32 multilayer perceptron (MLP) policy with two hidden layers ($[128, 64]$) in a massively parallel GPU simulation and enhance robustness by utilizing domain randomization over simulation parameters. We then study integer (Int8) quantization (per-tensor and per-feature) to allow for higher inference update rates on our resource-limited hardware, and we connect hardware power budgets to achievable update frequency via a cycles-per-update model for inference on our Cortex-M0. We propose a resource-aware gait scheduling viewpoint: given a device power budget, we can select the gait mode (trot/intermediate/gallop) that maximizes expected RL reward at a corresponding feasible update frequency. Finally, we deploy our MLP policy on a real-world large-scale robot on uneven terrain, qualitatively noting that domain-randomized training can improve out-of-distribution stability. We do not claim real-world large-robot empirical zero-shot transfer in this work.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24740
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Control of Microrobots with Reinforcement Learning under On-Device Compute Constraints
Liu, Yichen
Viswanadha, Kesava
Li, Zhongyu
Lojo, Nelson
Pister, Kristofer S. J.
Robotics
Systems and Control
An important function of autonomous microrobots is the ability to perform robust movement over terrain. This paper explores an edge ML approach to microrobot locomotion, allowing for on-device, lower latency control under compute, memory, and power constraints. This paper explores the locomotion of a sub-centimeter quadrupedal microrobot via reinforcement learning (RL) and deploys the resulting controller on an ultra-small system-on-chip (SoC), SC$μ$M-3C, featuring an ARM Cortex-M0 microcontroller running at 5 MHz. We train a compact FP32 multilayer perceptron (MLP) policy with two hidden layers ($[128, 64]$) in a massively parallel GPU simulation and enhance robustness by utilizing domain randomization over simulation parameters. We then study integer (Int8) quantization (per-tensor and per-feature) to allow for higher inference update rates on our resource-limited hardware, and we connect hardware power budgets to achievable update frequency via a cycles-per-update model for inference on our Cortex-M0. We propose a resource-aware gait scheduling viewpoint: given a device power budget, we can select the gait mode (trot/intermediate/gallop) that maximizes expected RL reward at a corresponding feasible update frequency. Finally, we deploy our MLP policy on a real-world large-scale robot on uneven terrain, qualitatively noting that domain-randomized training can improve out-of-distribution stability. We do not claim real-world large-robot empirical zero-shot transfer in this work.
title Control of Microrobots with Reinforcement Learning under On-Device Compute Constraints
topic Robotics
Systems and Control
url https://arxiv.org/abs/2512.24740