Design and Operation of Shared Machine Learning Clusters on Campus

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Kaiqiang, Sun, Decang, Wang, Hao, Ren, Zhenghang, Wan, Xinchen, Liao, Xudong, Wang, Zilong, Zhang, Junxue, Chen, Kai
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910942191681536
author Xu, Kaiqiang
Sun, Decang
Wang, Hao
Ren, Zhenghang
Wan, Xinchen
Liao, Xudong
Wang, Zilong
Zhang, Junxue
Chen, Kai
author_facet Xu, Kaiqiang
Sun, Decang
Wang, Hao
Ren, Zhenghang
Wan, Xinchen
Liao, Xudong
Wang, Zilong
Zhang, Junxue
Chen, Kai
contents Amid the rapid advancements in large machine learning (ML) models, universities worldwide are investing substantial funds and efforts into GPU clusters. However, managing a shared GPU cluster poses a pyramid of challenges, from hardware configuration to resource allocation among users. This paper introduces SING, a full-stack solution designed to streamline the management of shared GPU clusters in academic institutions. Motivated by the pressing need for efficient resource sharing and the challenges posed by limited staffing, we present a comprehensive view of SING's architecture and design choices, which achieves operational efficiency (i.e., low maintenance cost and high resource utilization). We also share experience and insights from the real-world operations of SING, including analysis of its usage patterns and management of incidents and failures. This paper is part of our ongoing effort to improve the management of shared ML clusters. We open-source relevant resources to facilitate the development and operation of similar clusters for ML.
format Preprint
id arxiv_https___arxiv_org_abs_2110_01556
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Design and Operation of Shared Machine Learning Clusters on Campus
Xu, Kaiqiang
Sun, Decang
Wang, Hao
Ren, Zhenghang
Wan, Xinchen
Liao, Xudong
Wang, Zilong
Zhang, Junxue
Chen, Kai
Distributed, Parallel, and Cluster Computing
Networking and Internet Architecture
Amid the rapid advancements in large machine learning (ML) models, universities worldwide are investing substantial funds and efforts into GPU clusters. However, managing a shared GPU cluster poses a pyramid of challenges, from hardware configuration to resource allocation among users. This paper introduces SING, a full-stack solution designed to streamline the management of shared GPU clusters in academic institutions. Motivated by the pressing need for efficient resource sharing and the challenges posed by limited staffing, we present a comprehensive view of SING's architecture and design choices, which achieves operational efficiency (i.e., low maintenance cost and high resource utilization). We also share experience and insights from the real-world operations of SING, including analysis of its usage patterns and management of incidents and failures. This paper is part of our ongoing effort to improve the management of shared ML clusters. We open-source relevant resources to facilitate the development and operation of similar clusters for ML.
title Design and Operation of Shared Machine Learning Clusters on Campus
topic Distributed, Parallel, and Cluster Computing
Networking and Internet Architecture
url https://arxiv.org/abs/2110.01556