Exploiting Stragglers in Distributed Computing Systems with Task Grouping

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Adikari, Tharindu, Al-Lawati, Haider, Lam, Jason, Hu, Zhenhua, Draper, Stark C.
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913572706058240
author Adikari, Tharindu
Al-Lawati, Haider
Lam, Jason
Hu, Zhenhua
Draper, Stark C.
author_facet Adikari, Tharindu
Al-Lawati, Haider
Lam, Jason
Hu, Zhenhua
Draper, Stark C.
contents We consider the problem of stragglers in distributed computing systems. Stragglers, which are compute nodes that unpredictably slow down, often increase the completion times of tasks. One common approach to mitigating stragglers is work replication, where only the first completion among replicated tasks is accepted, discarding the others. However, discarding work leads to resource wastage. In this paper, we propose a method for exploiting the work completed by stragglers rather than discarding it. The idea is to increase the granularity of the assigned work, and to increase the frequency of worker updates. We show that the proposed method reduces the completion time of tasks via experiments performed on a simulated cluster as well as on Amazon EC2 with Apache Hadoop.
format Preprint
id arxiv_https___arxiv_org_abs_2411_03645
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploiting Stragglers in Distributed Computing Systems with Task Grouping
Adikari, Tharindu
Al-Lawati, Haider
Lam, Jason
Hu, Zhenhua
Draper, Stark C.
Distributed, Parallel, and Cluster Computing
We consider the problem of stragglers in distributed computing systems. Stragglers, which are compute nodes that unpredictably slow down, often increase the completion times of tasks. One common approach to mitigating stragglers is work replication, where only the first completion among replicated tasks is accepted, discarding the others. However, discarding work leads to resource wastage. In this paper, we propose a method for exploiting the work completed by stragglers rather than discarding it. The idea is to increase the granularity of the assigned work, and to increase the frequency of worker updates. We show that the proposed method reduces the completion time of tasks via experiments performed on a simulated cluster as well as on Amazon EC2 with Apache Hadoop.
title Exploiting Stragglers in Distributed Computing Systems with Task Grouping
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2411.03645