Multi-Modal Decentralized Reinforcement Learning for Modular Reconfigurable Lunar Robots

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mishra, Ashutosh, Santra, Shreya, Neppel, Elian, Lombardi, Edoardo M. Rossi, Karimov, Shamistan, Uno, Kentaro, Yoshida, Kazuya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909865096511488
author Mishra, Ashutosh
Santra, Shreya
Neppel, Elian
Lombardi, Edoardo M. Rossi
Karimov, Shamistan
Uno, Kentaro
Yoshida, Kazuya
author_facet Mishra, Ashutosh
Santra, Shreya
Neppel, Elian
Lombardi, Edoardo M. Rossi
Karimov, Shamistan
Uno, Kentaro
Yoshida, Kazuya
contents Modular reconfigurable robots suit task-specific space operations, but the combinatorial growth of morphologies hinders unified control. We propose a decentralized reinforcement learning (Dec-RL) scheme where each module learns its own policy: wheel modules use Soft Actor-Critic (SAC) for locomotion and 7-DoF limbs use Proximal Policy Optimization (PPO) for steering and manipulation, enabling zero-shot generalization to unseen configurations. In simulation, the steering policy achieved a mean absolute error of 3.63° between desired and induced angles; the manipulation policy plateaued at 84.6 % success on a target-offset criterion; and the wheel policy cut average motor torque by 95.4 % relative to baseline while maintaining 99.6 % success. Lunar-analogue field tests validated zero-shot integration for autonomous locomotion, steering, and preliminary alignment for reconfiguration. The system transitioned smoothly among synchronous, parallel, and sequential modes for Policy Execution, without idle states or control conflicts, indicating a scalable, reusable, and robust approach for modular lunar robots.
format Preprint
id arxiv_https___arxiv_org_abs_2510_20347
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Modal Decentralized Reinforcement Learning for Modular Reconfigurable Lunar Robots
Mishra, Ashutosh
Santra, Shreya
Neppel, Elian
Lombardi, Edoardo M. Rossi
Karimov, Shamistan
Uno, Kentaro
Yoshida, Kazuya
Robotics
Multiagent Systems
Modular reconfigurable robots suit task-specific space operations, but the combinatorial growth of morphologies hinders unified control. We propose a decentralized reinforcement learning (Dec-RL) scheme where each module learns its own policy: wheel modules use Soft Actor-Critic (SAC) for locomotion and 7-DoF limbs use Proximal Policy Optimization (PPO) for steering and manipulation, enabling zero-shot generalization to unseen configurations. In simulation, the steering policy achieved a mean absolute error of 3.63° between desired and induced angles; the manipulation policy plateaued at 84.6 % success on a target-offset criterion; and the wheel policy cut average motor torque by 95.4 % relative to baseline while maintaining 99.6 % success. Lunar-analogue field tests validated zero-shot integration for autonomous locomotion, steering, and preliminary alignment for reconfiguration. The system transitioned smoothly among synchronous, parallel, and sequential modes for Policy Execution, without idle states or control conflicts, indicating a scalable, reusable, and robust approach for modular lunar robots.
title Multi-Modal Decentralized Reinforcement Learning for Modular Reconfigurable Lunar Robots
topic Robotics
Multiagent Systems
url https://arxiv.org/abs/2510.20347