Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective Elasticity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lv, Cunchi, Shi, Xiao, Lei, Zhengyu, Huang, Jinyue, Tan, Wenting, Zheng, Xiaohui, Zhao, Xiaofang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909529149538304
author Lv, Cunchi
Shi, Xiao
Lei, Zhengyu
Huang, Jinyue
Tan, Wenting
Zheng, Xiaohui
Zhao, Xiaofang
author_facet Lv, Cunchi
Shi, Xiao
Lei, Zhengyu
Huang, Jinyue
Tan, Wenting
Zheng, Xiaohui
Zhao, Xiaofang
contents Serverless computing, with its ease of management, auto-scaling, and cost-effectiveness, is widely adopted by deep learning (DL) applications. DL workloads, especially with large language models, require substantial GPU resources to ensure QoS. However, it is prone to produce GPU fragments (e.g., 15\%-94\%) in serverless DL systems due to the dynamicity of workloads and coarse-grained static GPU allocation mechanisms, gradually eroding the profits offered by serverless elasticity. Different from classical serverless systems that only scale horizontally, we present introspective elasticity (IE), a fine-grained and adaptive two-dimensional co-scaling mechanism to support GPU resourcing-on-demand for serverless DL tasks. Based on this insight, we build Dilu, a cross-layer and GPU-based serverless DL system with IE support. First, Dilu provides multi-factor profiling for DL tasks with efficient pruning search methods. Second, Dilu adheres to the resourcing-complementary principles in scheduling to improve GPU utilization with QoS guarantees. Third, Dilu adopts an adaptive 2D co-scaling method to enhance the elasticity of GPU provisioning in real time. Evaluations show that it can dynamically adjust the resourcing of various DL functions with low GPU fragmentation (10\%-46\% GPU defragmentation), high throughput (up to 1.8$\times$ inference and 1.1$\times$ training throughput increment) and QoS guarantees (11\%-71\% violation rate reduction), compared to the SOTA baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2503_05130
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective Elasticity
Lv, Cunchi
Shi, Xiao
Lei, Zhengyu
Huang, Jinyue
Tan, Wenting
Zheng, Xiaohui
Zhao, Xiaofang
Distributed, Parallel, and Cluster Computing
Serverless computing, with its ease of management, auto-scaling, and cost-effectiveness, is widely adopted by deep learning (DL) applications. DL workloads, especially with large language models, require substantial GPU resources to ensure QoS. However, it is prone to produce GPU fragments (e.g., 15\%-94\%) in serverless DL systems due to the dynamicity of workloads and coarse-grained static GPU allocation mechanisms, gradually eroding the profits offered by serverless elasticity. Different from classical serverless systems that only scale horizontally, we present introspective elasticity (IE), a fine-grained and adaptive two-dimensional co-scaling mechanism to support GPU resourcing-on-demand for serverless DL tasks. Based on this insight, we build Dilu, a cross-layer and GPU-based serverless DL system with IE support. First, Dilu provides multi-factor profiling for DL tasks with efficient pruning search methods. Second, Dilu adheres to the resourcing-complementary principles in scheduling to improve GPU utilization with QoS guarantees. Third, Dilu adopts an adaptive 2D co-scaling method to enhance the elasticity of GPU provisioning in real time. Evaluations show that it can dynamically adjust the resourcing of various DL functions with low GPU fragmentation (10\%-46\% GPU defragmentation), high throughput (up to 1.8$\times$ inference and 1.1$\times$ training throughput increment) and QoS guarantees (11\%-71\% violation rate reduction), compared to the SOTA baselines.
title Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective Elasticity
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2503.05130