Design and Operation of Shared Machine Learning Clusters on Campus
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2021
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910942191681536 |
|---|---|
| author | Xu, Kaiqiang Sun, Decang Wang, Hao Ren, Zhenghang Wan, Xinchen Liao, Xudong Wang, Zilong Zhang, Junxue Chen, Kai |
| author_facet | Xu, Kaiqiang Sun, Decang Wang, Hao Ren, Zhenghang Wan, Xinchen Liao, Xudong Wang, Zilong Zhang, Junxue Chen, Kai |
| contents | Amid the rapid advancements in large machine learning (ML) models, universities worldwide are investing substantial funds and efforts into GPU clusters. However, managing a shared GPU cluster poses a pyramid of challenges, from hardware configuration to resource allocation among users.
This paper introduces SING, a full-stack solution designed to streamline the management of shared GPU clusters in academic institutions. Motivated by the pressing need for efficient resource sharing and the challenges posed by limited staffing, we present a comprehensive view of SING's architecture and design choices, which achieves operational efficiency (i.e., low maintenance cost and high resource utilization). We also share experience and insights from the real-world operations of SING, including analysis of its usage patterns and management of incidents and failures.
This paper is part of our ongoing effort to improve the management of shared ML clusters. We open-source relevant resources to facilitate the development and operation of similar clusters for ML. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2110_01556 |
| institution | arXiv |
| publishDate | 2021 |
| record_format | arxiv |
| spellingShingle | Design and Operation of Shared Machine Learning Clusters on Campus Xu, Kaiqiang Sun, Decang Wang, Hao Ren, Zhenghang Wan, Xinchen Liao, Xudong Wang, Zilong Zhang, Junxue Chen, Kai Distributed, Parallel, and Cluster Computing Networking and Internet Architecture Amid the rapid advancements in large machine learning (ML) models, universities worldwide are investing substantial funds and efforts into GPU clusters. However, managing a shared GPU cluster poses a pyramid of challenges, from hardware configuration to resource allocation among users. This paper introduces SING, a full-stack solution designed to streamline the management of shared GPU clusters in academic institutions. Motivated by the pressing need for efficient resource sharing and the challenges posed by limited staffing, we present a comprehensive view of SING's architecture and design choices, which achieves operational efficiency (i.e., low maintenance cost and high resource utilization). We also share experience and insights from the real-world operations of SING, including analysis of its usage patterns and management of incidents and failures. This paper is part of our ongoing effort to improve the management of shared ML clusters. We open-source relevant resources to facilitate the development and operation of similar clusters for ML. |
| title | Design and Operation of Shared Machine Learning Clusters on Campus |
| topic | Distributed, Parallel, and Cluster Computing Networking and Internet Architecture |
| url | https://arxiv.org/abs/2110.01556 |