Compass: A Decentralized Scheduler for Latency-Sensitive ML Workflows

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Yuting, Merlina, Andrea, Song, Weijia, Yuan, Tiancheng, Birman, Ken, Vitenberg, Roman
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910347167793152
author Yang, Yuting
Merlina, Andrea
Song, Weijia
Yuan, Tiancheng
Birman, Ken
Vitenberg, Roman
author_facet Yang, Yuting
Merlina, Andrea
Song, Weijia
Yuan, Tiancheng
Birman, Ken
Vitenberg, Roman
contents We consider ML query processing in distributed systems where GPU-enabled workers coordinate to execute complex queries: a computing style often seen in applications that interact with users in support of image processing and natural language processing. In such systems, coscheduling of GPU memory management and task placement represents a promising opportunity. We propose Compass, a novel framework that unifies these functions to reduce job latency while using resources efficiently, placing tasks where data dependencies will be satisfied, collocating tasks from the same job (when this will not overload the host or its GPU), and efficiently managing GPU memory. Comparison with other state of the art schedulers shows a significant reduction in completion times while requiring the same amount or even fewer resources. In one case, just half the servers were needed for processing the same workload.
format Preprint
id arxiv_https___arxiv_org_abs_2402_17652
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Compass: A Decentralized Scheduler for Latency-Sensitive ML Workflows
Yang, Yuting
Merlina, Andrea
Song, Weijia
Yuan, Tiancheng
Birman, Ken
Vitenberg, Roman
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
We consider ML query processing in distributed systems where GPU-enabled workers coordinate to execute complex queries: a computing style often seen in applications that interact with users in support of image processing and natural language processing. In such systems, coscheduling of GPU memory management and task placement represents a promising opportunity. We propose Compass, a novel framework that unifies these functions to reduce job latency while using resources efficiently, placing tasks where data dependencies will be satisfied, collocating tasks from the same job (when this will not overload the host or its GPU), and efficiently managing GPU memory. Comparison with other state of the art schedulers shows a significant reduction in completion times while requiring the same amount or even fewer resources. In one case, just half the servers were needed for processing the same workload.
title Compass: A Decentralized Scheduler for Latency-Sensitive ML Workflows
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2402.17652