PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yeo, Gwangoo, Kim, Jiin, Choi, Yujeong, Rhu, Minsoo
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910719785566208
author Yeo, Gwangoo
Kim, Jiin
Choi, Yujeong
Rhu, Minsoo
author_facet Yeo, Gwangoo
Kim, Jiin
Choi, Yujeong
Rhu, Minsoo
contents NVIDIA's Multi-Instance GPU (MIG) is a feature that enables system designers to reconfigure one large GPU into multiple smaller GPU slices. This work characterizes this emerging GPU and evaluates its effectiveness in designing high-performance AI inference servers. Our study reveals that the data preprocessing stage of AI inference causes significant performance bottlenecks to MIG. To this end, we present PREBA, which is a hardware/software co-design targeting MIG inference servers. Our first proposition is an FPGA-based data preprocessing accelerator that unlocks the full potential of MIG with domain-specific acceleration of data preprocessing. The MIG inference server unleashed from preprocessing overheads is then augmented with our dynamic batching system that enables high-performance inference. PREBA is implemented end-to-end in real systems, providing a 3.7x improvement in throughput, 3.4x reduction in tail latency, 3.5x improvement in energy-efficiency, and 3.0x improvement in cost-efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2411_19114
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
Yeo, Gwangoo
Kim, Jiin
Choi, Yujeong
Rhu, Minsoo
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Hardware Architecture
Machine Learning
NVIDIA's Multi-Instance GPU (MIG) is a feature that enables system designers to reconfigure one large GPU into multiple smaller GPU slices. This work characterizes this emerging GPU and evaluates its effectiveness in designing high-performance AI inference servers. Our study reveals that the data preprocessing stage of AI inference causes significant performance bottlenecks to MIG. To this end, we present PREBA, which is a hardware/software co-design targeting MIG inference servers. Our first proposition is an FPGA-based data preprocessing accelerator that unlocks the full potential of MIG with domain-specific acceleration of data preprocessing. The MIG inference server unleashed from preprocessing overheads is then augmented with our dynamic batching system that enables high-performance inference. PREBA is implemented end-to-end in real systems, providing a 3.7x improvement in throughput, 3.4x reduction in tail latency, 3.5x improvement in energy-efficiency, and 3.0x improvement in cost-efficiency.
title PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Hardware Architecture
Machine Learning
url https://arxiv.org/abs/2411.19114