Multi-Agent Reinforcement Learning Counteracts Delayed CSI in Multi-Satellite Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aristodemou, Marios, Omid, Yasaman, Lambotharan, Sangarapillai, Derakhshan, Mahsa, Hanzo, Lajos
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912970967089152
author Aristodemou, Marios
Omid, Yasaman
Lambotharan, Sangarapillai
Derakhshan, Mahsa
Hanzo, Lajos
author_facet Aristodemou, Marios
Omid, Yasaman
Lambotharan, Sangarapillai
Derakhshan, Mahsa
Hanzo, Lajos
contents The integration of satellite communication networks with next-generation (NG) technologies is a promising approach towards global connectivity. However, the quality of services is highly dependant on the availability of accurate channel state information (CSI). Channel estimation in satellite communications is challenging due to the high propagation delay between terrestrial users and satellites, which results in outdated CSI observations on the satellite side. In this paper, we study the downlink transmission of multiple satellites acting as distributed base stations (BS) to mobile terrestrial users. We propose a multi-agent reinforcement learning (MARL) algorithm which aims for maximising the sum-rate of the users, while coping with the outdated CSI. We design a novel bi-level optimisation, procedure themes as dual stage proximal policy optimisation (DS-PPO), for tackling the problem of large continuous action spaces as well as of independent and non-identically distributed (non-IID) environments in MARL. Specifically, the first stage of DS-PPO maximises the sum-rate for an individual satellite and the second stage maximises the sum-rate when all the satellites cooperate to form a distributed multi-antenna BS. Our numerical results demonstrate the robustness of DS-PPO to CSI imperfections as well as the sum-rate improvement attached by the use of DS-PPO. In addition, we provide the convergence analysis for the DS-PPO along with the computational complexity.
format Preprint
id arxiv_https___arxiv_org_abs_2603_16470
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multi-Agent Reinforcement Learning Counteracts Delayed CSI in Multi-Satellite Systems
Aristodemou, Marios
Omid, Yasaman
Lambotharan, Sangarapillai
Derakhshan, Mahsa
Hanzo, Lajos
Information Theory
Artificial Intelligence
Signal Processing
The integration of satellite communication networks with next-generation (NG) technologies is a promising approach towards global connectivity. However, the quality of services is highly dependant on the availability of accurate channel state information (CSI). Channel estimation in satellite communications is challenging due to the high propagation delay between terrestrial users and satellites, which results in outdated CSI observations on the satellite side. In this paper, we study the downlink transmission of multiple satellites acting as distributed base stations (BS) to mobile terrestrial users. We propose a multi-agent reinforcement learning (MARL) algorithm which aims for maximising the sum-rate of the users, while coping with the outdated CSI. We design a novel bi-level optimisation, procedure themes as dual stage proximal policy optimisation (DS-PPO), for tackling the problem of large continuous action spaces as well as of independent and non-identically distributed (non-IID) environments in MARL. Specifically, the first stage of DS-PPO maximises the sum-rate for an individual satellite and the second stage maximises the sum-rate when all the satellites cooperate to form a distributed multi-antenna BS. Our numerical results demonstrate the robustness of DS-PPO to CSI imperfections as well as the sum-rate improvement attached by the use of DS-PPO. In addition, we provide the convergence analysis for the DS-PPO along with the computational complexity.
title Multi-Agent Reinforcement Learning Counteracts Delayed CSI in Multi-Satellite Systems
topic Information Theory
Artificial Intelligence
Signal Processing
url https://arxiv.org/abs/2603.16470