EdgeSight: Enabling Modeless and Cost-Efficient Inference at the Edge

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lao, ChonLam, Gao, Jiaqi, Ananthanarayanan, Ganesh, Akella, Aditya, Yu, Minlan
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909456766337024
author Lao, ChonLam
Gao, Jiaqi
Ananthanarayanan, Ganesh
Akella, Aditya
Yu, Minlan
author_facet Lao, ChonLam
Gao, Jiaqi
Ananthanarayanan, Ganesh
Akella, Aditya
Yu, Minlan
contents Traditional ML inference is evolving toward modeless inference, which abstracts the complexity of model selection from users, allowing the system to automatically choose the most appropriate model for each request based on accuracy and resource requirements. While prior studies have focused on modeless inference within data centers, this paper tackles the pressing need for cost-efficient modeless inference at the edge -- particularly within its unique constraints of limited device memory, volatile network conditions, and restricted power consumption. To overcome these challenges, we propose EdgeSight, a system that provides cost-efficient EdgeSight serving for diverse DNNs at the edge. EdgeSight employs an edge-data center (edge-DC) architecture, utilizing confidence scaling to reduce the number of model options while meeting diverse accuracy requirements. Additionally, it supports lossy inference in volatile network environments. Our experimental results show that EdgeSight outperforms existing systems by up to 1.6x in P99 latency for modeless services. Furthermore, our FPGA prototype demonstrates similar performance at certain accuracy levels, with a power consumption reduction of up to 3.34x.
format Preprint
id arxiv_https___arxiv_org_abs_2405_19213
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EdgeSight: Enabling Modeless and Cost-Efficient Inference at the Edge
Lao, ChonLam
Gao, Jiaqi
Ananthanarayanan, Ganesh
Akella, Aditya
Yu, Minlan
Systems and Control
Artificial Intelligence
Machine Learning
Networking and Internet Architecture
Traditional ML inference is evolving toward modeless inference, which abstracts the complexity of model selection from users, allowing the system to automatically choose the most appropriate model for each request based on accuracy and resource requirements. While prior studies have focused on modeless inference within data centers, this paper tackles the pressing need for cost-efficient modeless inference at the edge -- particularly within its unique constraints of limited device memory, volatile network conditions, and restricted power consumption. To overcome these challenges, we propose EdgeSight, a system that provides cost-efficient EdgeSight serving for diverse DNNs at the edge. EdgeSight employs an edge-data center (edge-DC) architecture, utilizing confidence scaling to reduce the number of model options while meeting diverse accuracy requirements. Additionally, it supports lossy inference in volatile network environments. Our experimental results show that EdgeSight outperforms existing systems by up to 1.6x in P99 latency for modeless services. Furthermore, our FPGA prototype demonstrates similar performance at certain accuracy levels, with a power consumption reduction of up to 3.34x.
title EdgeSight: Enabling Modeless and Cost-Efficient Inference at the Edge
topic Systems and Control
Artificial Intelligence
Machine Learning
Networking and Internet Architecture
url https://arxiv.org/abs/2405.19213