FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge Environments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Li, Hanafy, Walid A., Abdelzaher, Tarek, Irwin, David, Milzman, Jesse, Shenoy, Prashant
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917097659957248
author Wu, Li
Hanafy, Walid A.
Abdelzaher, Tarek
Irwin, David
Milzman, Jesse
Shenoy, Prashant
author_facet Wu, Li
Hanafy, Walid A.
Abdelzaher, Tarek
Irwin, David
Milzman, Jesse
Shenoy, Prashant
contents Model serving systems have become popular for deploying deep learning models for various latency-sensitive inference tasks. While traditional replication-based methods have been used for failure-resilient model serving in the cloud, such methods are often infeasible in edge environments due to significant resource constraints that preclude full replication. To address this problem, this paper presents FailLite, a failure-resilient model serving system that employs (i) a heterogeneous replication where failover models are smaller variants of the original model, (ii) an intelligent approach that uses warm replicas to ensure quick failover for critical applications while using cold replicas, and (iii) progressive failover to provide low mean time to recovery (MTTR) for the remaining applications. We implement a full prototype of our system and demonstrate its efficacy on an experimental edge testbed. Our results using 27 models show that FailLite can recover all failed applications with 175.5ms MTTR and only a 0.6% reduction in accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15856
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge Environments
Wu, Li
Hanafy, Walid A.
Abdelzaher, Tarek
Irwin, David
Milzman, Jesse
Shenoy, Prashant
Distributed, Parallel, and Cluster Computing
Model serving systems have become popular for deploying deep learning models for various latency-sensitive inference tasks. While traditional replication-based methods have been used for failure-resilient model serving in the cloud, such methods are often infeasible in edge environments due to significant resource constraints that preclude full replication. To address this problem, this paper presents FailLite, a failure-resilient model serving system that employs (i) a heterogeneous replication where failover models are smaller variants of the original model, (ii) an intelligent approach that uses warm replicas to ensure quick failover for critical applications while using cold replicas, and (iii) progressive failover to provide low mean time to recovery (MTTR) for the remaining applications. We implement a full prototype of our system and demonstrate its efficacy on an experimental edge testbed. Our results using 27 models show that FailLite can recover all failed applications with 175.5ms MTTR and only a 0.6% reduction in accuracy.
title FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge Environments
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2504.15856