AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Abhishek Vijaya, Antichi, Gianni, Singh, Rachee
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912239529754624
author Kumar, Abhishek Vijaya
Antichi, Gianni
Singh, Rachee
author_facet Kumar, Abhishek Vijaya
Antichi, Gianni
Singh, Rachee
contents Inference on large-language models (LLMs) is constrained by GPU memory capacity. A sudden increase in the number of inference requests to a cloud-hosted LLM can deplete GPU memory, leading to contention between multiple prompts for limited resources. Modern LLM serving engines deal with the challenge of limited GPU memory using admission control, which causes them to be unresponsive during request bursts. We propose that preemptive scheduling of prompts in time slices is essential for ensuring responsive LLM inference, especially under conditions of high load and limited GPU memory. However, preempting prompt inference incurs a high paging overhead, which reduces inference throughput. We present Aqua, a GPU memory management framework that significantly reduces the overhead of paging inference state, achieving both responsive and high-throughput inference even under bursty request patterns. We evaluate Aqua by hosting several state-of-the-art large generative ML models of different modalities on servers with 8 Nvidia H100 80G GPUs. Aqua improves the responsiveness of LLM inference by 20X compared to the state-of-the-art and improves LLM inference throughput over a single long prompt by 4X.
format Preprint
id arxiv_https___arxiv_org_abs_2407_21255
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
Kumar, Abhishek Vijaya
Antichi, Gianni
Singh, Rachee
Distributed, Parallel, and Cluster Computing
Inference on large-language models (LLMs) is constrained by GPU memory capacity. A sudden increase in the number of inference requests to a cloud-hosted LLM can deplete GPU memory, leading to contention between multiple prompts for limited resources. Modern LLM serving engines deal with the challenge of limited GPU memory using admission control, which causes them to be unresponsive during request bursts. We propose that preemptive scheduling of prompts in time slices is essential for ensuring responsive LLM inference, especially under conditions of high load and limited GPU memory. However, preempting prompt inference incurs a high paging overhead, which reduces inference throughput. We present Aqua, a GPU memory management framework that significantly reduces the overhead of paging inference state, achieving both responsive and high-throughput inference even under bursty request patterns. We evaluate Aqua by hosting several state-of-the-art large generative ML models of different modalities on servers with 8 Nvidia H100 80G GPUs. Aqua improves the responsiveness of LLM inference by 20X compared to the state-of-the-art and improves LLM inference throughput over a single long prompt by 4X.
title AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2407.21255