FIRE: A Failure-Adaptive Reinforcement Learning Framework for Edge Computing Migrations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Siew, Marie, Sharma, Shikhar, Li, Zekai, Guo, Kun, Xu, Chao, Lorido-Botran, Tania, Quek, Tony Q. S., Joe-Wong, Carlee
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915561665986560
author Siew, Marie
Sharma, Shikhar
Li, Zekai
Guo, Kun
Xu, Chao
Lorido-Botran, Tania
Quek, Tony Q. S.
Joe-Wong, Carlee
author_facet Siew, Marie
Sharma, Shikhar
Li, Zekai
Guo, Kun
Xu, Chao
Lorido-Botran, Tania
Quek, Tony Q. S.
Joe-Wong, Carlee
contents In edge computing, users' service profiles are migrated due to user mobility. Reinforcement learning (RL) frameworks have been proposed to do so, often trained on simulated data. However, existing RL frameworks overlook occasional server failures, which although rare, impact latency-sensitive applications like autonomous driving and real-time obstacle detection. Nevertheless, these failures (rare events), being not adequately represented in historical training data, pose a challenge for data-driven RL algorithms. As it is impractical to adjust failure frequency in real-world applications for training, we introduce FIRE, a framework that adapts to rare events by training a RL policy in an edge computing digital twin environment. We propose ImRE, an importance sampling-based Q-learning algorithm, which samples rare events proportionally to their impact on the value function. FIRE considers delay, migration, failure, and backup placement costs across individual and shared service profiles. We prove ImRE's boundedness and convergence to optimality. Next, we introduce novel deep Q-learning (ImDQL) and actor critic (ImACRE) versions of our algorithm to enhance scalability. We extend our framework to accommodate users with varying risk tolerances. Through trace driven experiments, we show that FIRE reduces costs compared to vanilla RL and the greedy baseline in the event of failures.
format Preprint
id arxiv_https___arxiv_org_abs_2209_14399
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle FIRE: A Failure-Adaptive Reinforcement Learning Framework for Edge Computing Migrations
Siew, Marie
Sharma, Shikhar
Li, Zekai
Guo, Kun
Xu, Chao
Lorido-Botran, Tania
Quek, Tony Q. S.
Joe-Wong, Carlee
Networking and Internet Architecture
Machine Learning
Systems and Control
In edge computing, users' service profiles are migrated due to user mobility. Reinforcement learning (RL) frameworks have been proposed to do so, often trained on simulated data. However, existing RL frameworks overlook occasional server failures, which although rare, impact latency-sensitive applications like autonomous driving and real-time obstacle detection. Nevertheless, these failures (rare events), being not adequately represented in historical training data, pose a challenge for data-driven RL algorithms. As it is impractical to adjust failure frequency in real-world applications for training, we introduce FIRE, a framework that adapts to rare events by training a RL policy in an edge computing digital twin environment. We propose ImRE, an importance sampling-based Q-learning algorithm, which samples rare events proportionally to their impact on the value function. FIRE considers delay, migration, failure, and backup placement costs across individual and shared service profiles. We prove ImRE's boundedness and convergence to optimality. Next, we introduce novel deep Q-learning (ImDQL) and actor critic (ImACRE) versions of our algorithm to enhance scalability. We extend our framework to accommodate users with varying risk tolerances. Through trace driven experiments, we show that FIRE reduces costs compared to vanilla RL and the greedy baseline in the event of failures.
title FIRE: A Failure-Adaptive Reinforcement Learning Framework for Edge Computing Migrations
topic Networking and Internet Architecture
Machine Learning
Systems and Control
url https://arxiv.org/abs/2209.14399