Scheduling Data-Intensive Workloads in Large-Scale Distributed Systems: Trends and Challenges

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Stavrinides, Georgios L., Karatza, Helen D.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918177789706240
author Stavrinides, Georgios L.
Karatza, Helen D.
author_facet Stavrinides, Georgios L.
Karatza, Helen D.
contents With the explosive growth of big data, workloads tend to get more complex and computationally demanding. Such applications are processed on distributed interconnected resources that are becoming larger in scale and computational capacity. Data-intensive applications may have different degrees of parallelism and must effectively exploit data locality. Furthermore, they may impose several Quality of Service requirements, such as time constraints and resilience against failures, as well as other objectives, like energy efficiency. These features of the workloads, as well as the inherent characteristics of the computing resources required to process them, present major challenges that require the employment of effective scheduling techniques. In this chapter, a classification of data-intensive workloads is proposed and an overview of the most commonly used approaches for their scheduling in large-scale distributed systems is given. We present novel strategies that have been proposed in the literature and shed light on open challenges and future directions.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25362
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scheduling Data-Intensive Workloads in Large-Scale Distributed Systems: Trends and Challenges
Stavrinides, Georgios L.
Karatza, Helen D.
Distributed, Parallel, and Cluster Computing
With the explosive growth of big data, workloads tend to get more complex and computationally demanding. Such applications are processed on distributed interconnected resources that are becoming larger in scale and computational capacity. Data-intensive applications may have different degrees of parallelism and must effectively exploit data locality. Furthermore, they may impose several Quality of Service requirements, such as time constraints and resilience against failures, as well as other objectives, like energy efficiency. These features of the workloads, as well as the inherent characteristics of the computing resources required to process them, present major challenges that require the employment of effective scheduling techniques. In this chapter, a classification of data-intensive workloads is proposed and an overview of the most commonly used approaches for their scheduling in large-scale distributed systems is given. We present novel strategies that have been proposed in the literature and shed light on open challenges and future directions.
title Scheduling Data-Intensive Workloads in Large-Scale Distributed Systems: Trends and Challenges
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2510.25362