Dispatching Odyssey: Exploring Performance in Computing Clusters under Real-world Workloads

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yildiz, Mert, Rolich, Alexey, Baiocchi, Andrea
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914504009318400
author Yildiz, Mert
Rolich, Alexey
Baiocchi, Andrea
author_facet Yildiz, Mert
Rolich, Alexey
Baiocchi, Andrea
contents Recent workload measurements in Google data centers provide an opportunity to challenge existing models and, more broadly, to enhance the understanding of dispatching policies in computing clusters. Through extensive data-driven simulations, we aim to highlight the key features of workload traffic traces that influence response time performance under simple yet representative dispatching policies. For a given computational power budget, we vary the cluster size, i.e., the number of available servers. A job-level analysis reveals that Join Idle Queue (JIQ) and Least Work Left (LWL) exhibit an optimal working point for a fixed utilization coefficient as the number of servers is varied, whereas Round Robin (RR) demonstrates monotonously worsening performance. Additionally, we explore the accuracy of simple G/G queue approximations. When decomposing jobs into tasks, interesting results emerge; notably, the simpler, non-size-based policy JIQ appears to outperform the more "powerful" size-based LWL policy. Complementing these findings, we present preliminary results on a two-stage scheduling approach that partitions tasks based on service thresholds, illustrating that modest architectural modifications can further enhance performance under realistic workload conditions. We provide insights into these results and suggest promising directions for fully explaining the observed phenomena.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10184
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dispatching Odyssey: Exploring Performance in Computing Clusters under Real-world Workloads
Yildiz, Mert
Rolich, Alexey
Baiocchi, Andrea
Distributed, Parallel, and Cluster Computing
Recent workload measurements in Google data centers provide an opportunity to challenge existing models and, more broadly, to enhance the understanding of dispatching policies in computing clusters. Through extensive data-driven simulations, we aim to highlight the key features of workload traffic traces that influence response time performance under simple yet representative dispatching policies. For a given computational power budget, we vary the cluster size, i.e., the number of available servers. A job-level analysis reveals that Join Idle Queue (JIQ) and Least Work Left (LWL) exhibit an optimal working point for a fixed utilization coefficient as the number of servers is varied, whereas Round Robin (RR) demonstrates monotonously worsening performance. Additionally, we explore the accuracy of simple G/G queue approximations. When decomposing jobs into tasks, interesting results emerge; notably, the simpler, non-size-based policy JIQ appears to outperform the more "powerful" size-based LWL policy. Complementing these findings, we present preliminary results on a two-stage scheduling approach that partitions tasks based on service thresholds, illustrating that modest architectural modifications can further enhance performance under realistic workload conditions. We provide insights into these results and suggest promising directions for fully explaining the observed phenomena.
title Dispatching Odyssey: Exploring Performance in Computing Clusters under Real-world Workloads
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2504.10184