Multi-Satellite Beam Hopping and Power Allocation Using Deep Reinforcement Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xie, Xia, Fan, Kexin, Deng, Wenfeng, Pappas, Nikolaos, Zhang, Qinyu
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912176818618368
author Xie, Xia
Fan, Kexin
Deng, Wenfeng
Pappas, Nikolaos
Zhang, Qinyu
author_facet Xie, Xia
Fan, Kexin
Deng, Wenfeng
Pappas, Nikolaos
Zhang, Qinyu
contents In non-geostationary orbit (NGSO) satellite communication systems, effectively utilizing beam hopping (BH) technology is crucial for addressing uneven traffic demands. However, optimizing beam scheduling and resource allocation in multi-NGSO BH scenarios remains a significant challenge. This paper proposes a multi-NGSO BH algorithm based on deep reinforcement learning (DRL) to optimize beam illumination patterns and power allocation. By leveraging three degrees of freedom (i.e., time, space, and power), the algorithm aims to optimize the long-term throughput and the long-term cumulative average delay (LTCAD). The solution is based on proximal policy optimization (PPO) with a hybrid action space combining discrete and continuous actions. Using two policy networks with a shared base layer, the proposed algorithm jointly optimizes beam scheduling and power allocation. One network selects beam illumination patterns in the discrete action space, while the other manages power allocation in the continuous space. Simulation results show that the proposed algorithm significantly reduces LTCAD while maintaining high throughput in time-varying traffic scenarios. Compared to the four benchmark methods, it improves network throughput by up to $8.9\%$ and reduces LTCAD by up to $69.2\%$
format Preprint
id arxiv_https___arxiv_org_abs_2501_02309
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Satellite Beam Hopping and Power Allocation Using Deep Reinforcement Learning
Xie, Xia
Fan, Kexin
Deng, Wenfeng
Pappas, Nikolaos
Zhang, Qinyu
Systems and Control
In non-geostationary orbit (NGSO) satellite communication systems, effectively utilizing beam hopping (BH) technology is crucial for addressing uneven traffic demands. However, optimizing beam scheduling and resource allocation in multi-NGSO BH scenarios remains a significant challenge. This paper proposes a multi-NGSO BH algorithm based on deep reinforcement learning (DRL) to optimize beam illumination patterns and power allocation. By leveraging three degrees of freedom (i.e., time, space, and power), the algorithm aims to optimize the long-term throughput and the long-term cumulative average delay (LTCAD). The solution is based on proximal policy optimization (PPO) with a hybrid action space combining discrete and continuous actions. Using two policy networks with a shared base layer, the proposed algorithm jointly optimizes beam scheduling and power allocation. One network selects beam illumination patterns in the discrete action space, while the other manages power allocation in the continuous space. Simulation results show that the proposed algorithm significantly reduces LTCAD while maintaining high throughput in time-varying traffic scenarios. Compared to the four benchmark methods, it improves network throughput by up to $8.9\%$ and reduces LTCAD by up to $69.2\%$
title Multi-Satellite Beam Hopping and Power Allocation Using Deep Reinforcement Learning
topic Systems and Control
url https://arxiv.org/abs/2501.02309