PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866910719785566208 |
|---|---|
| author | Yeo, Gwangoo Kim, Jiin Choi, Yujeong Rhu, Minsoo |
| author_facet | Yeo, Gwangoo Kim, Jiin Choi, Yujeong Rhu, Minsoo |
| contents | NVIDIA's Multi-Instance GPU (MIG) is a feature that enables system designers to reconfigure one large GPU into multiple smaller GPU slices. This work characterizes this emerging GPU and evaluates its effectiveness in designing high-performance AI inference servers. Our study reveals that the data preprocessing stage of AI inference causes significant performance bottlenecks to MIG. To this end, we present PREBA, which is a hardware/software co-design targeting MIG inference servers. Our first proposition is an FPGA-based data preprocessing accelerator that unlocks the full potential of MIG with domain-specific acceleration of data preprocessing. The MIG inference server unleashed from preprocessing overheads is then augmented with our dynamic batching system that enables high-performance inference. PREBA is implemented end-to-end in real systems, providing a 3.7x improvement in throughput, 3.4x reduction in tail latency, 3.5x improvement in energy-efficiency, and 3.0x improvement in cost-efficiency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_19114 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers Yeo, Gwangoo Kim, Jiin Choi, Yujeong Rhu, Minsoo Distributed, Parallel, and Cluster Computing Artificial Intelligence Hardware Architecture Machine Learning NVIDIA's Multi-Instance GPU (MIG) is a feature that enables system designers to reconfigure one large GPU into multiple smaller GPU slices. This work characterizes this emerging GPU and evaluates its effectiveness in designing high-performance AI inference servers. Our study reveals that the data preprocessing stage of AI inference causes significant performance bottlenecks to MIG. To this end, we present PREBA, which is a hardware/software co-design targeting MIG inference servers. Our first proposition is an FPGA-based data preprocessing accelerator that unlocks the full potential of MIG with domain-specific acceleration of data preprocessing. The MIG inference server unleashed from preprocessing overheads is then augmented with our dynamic batching system that enables high-performance inference. PREBA is implemented end-to-end in real systems, providing a 3.7x improvement in throughput, 3.4x reduction in tail latency, 3.5x improvement in energy-efficiency, and 3.0x improvement in cost-efficiency. |
| title | PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers |
| topic | Distributed, Parallel, and Cluster Computing Artificial Intelligence Hardware Architecture Machine Learning |
| url | https://arxiv.org/abs/2411.19114 |