Machine Learning-Driven Predictive Resource Management in Complex Science Workflows

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chowdhury, Tasnuva, Maeno, Tadashi, Akman, Fatih Furkan, Boudreau, Joseph, Dutta, Sankha, Feng, Shengyu, Hoisie, Adolfy, Hsu, Kuan-Chieh, Khan, Raees, Kim, Jaehyung, Kilic, Ozgur O., Klasky, Scott, Klimentov, Alexei, Korchuganova, Tatiana, Outschoorn, Verena Ingrid Martinez, Nilsson, Paul, Park, David K., Podhorszki, Norbert, Ren, Yihui, Steele, John Rembrandt, Suter, Frédéric, Vatsavai, Sairam Sri, Wenaus, Torre, Yang, Wei, Yang, Yiming, Yoo, Shinjae
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915686728597504
author Chowdhury, Tasnuva
Maeno, Tadashi
Akman, Fatih Furkan
Boudreau, Joseph
Dutta, Sankha
Feng, Shengyu
Hoisie, Adolfy
Hsu, Kuan-Chieh
Khan, Raees
Kim, Jaehyung
Kilic, Ozgur O.
Klasky, Scott
Klimentov, Alexei
Korchuganova, Tatiana
Outschoorn, Verena Ingrid Martinez
Nilsson, Paul
Park, David K.
Podhorszki, Norbert
Ren, Yihui
Steele, John Rembrandt
Suter, Frédéric
Vatsavai, Sairam Sri
Wenaus, Torre
Yang, Wei
Yang, Yiming
Yoo, Shinjae
author_facet Chowdhury, Tasnuva
Maeno, Tadashi
Akman, Fatih Furkan
Boudreau, Joseph
Dutta, Sankha
Feng, Shengyu
Hoisie, Adolfy
Hsu, Kuan-Chieh
Khan, Raees
Kim, Jaehyung
Kilic, Ozgur O.
Klasky, Scott
Klimentov, Alexei
Korchuganova, Tatiana
Outschoorn, Verena Ingrid Martinez
Nilsson, Paul
Park, David K.
Podhorszki, Norbert
Ren, Yihui
Steele, John Rembrandt
Suter, Frédéric
Vatsavai, Sairam Sri
Wenaus, Torre
Yang, Wei
Yang, Yiming
Yoo, Shinjae
contents The collaborative efforts of large communities in science experiments, often comprising thousands of global members, reflect a monumental commitment to exploration and discovery. Recently, advanced and complex data processing has gained increasing importance in science experiments. Data processing workflows typically consist of multiple intricate steps, and the precise specification of resource requirements is crucial for each step to allocate optimal resources for effective processing. Estimating resource requirements in advance is challenging due to a wide range of analysis scenarios, varying skill levels among community members, and the continuously increasing spectrum of computing options. One practical approach to mitigate these challenges involves initially processing a subset of each step to measure precise resource utilization from actual processing profiles before completing the entire step. While this two-staged approach enables processing on optimal resources for most of the workflow, it has drawbacks such as initial inaccuracies leading to potential failures and suboptimal resource usage, along with overhead from waiting for initial processing completion, which is critical for fast-turnaround analyses. In this context, our study introduces a novel pipeline of machine learning models within a comprehensive workflow management system, the Production and Distributed Analysis (PanDA) system. These models employ advanced machine learning techniques to predict key resource requirements, overcoming challenges posed by limited upfront knowledge of characteristics at each step. Accurate forecasts of resource requirements enable informed and proactive decision-making in workflow management, enhancing the efficiency of handling diverse, complex workflows across heterogeneous resources.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11512
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Machine Learning-Driven Predictive Resource Management in Complex Science Workflows
Chowdhury, Tasnuva
Maeno, Tadashi
Akman, Fatih Furkan
Boudreau, Joseph
Dutta, Sankha
Feng, Shengyu
Hoisie, Adolfy
Hsu, Kuan-Chieh
Khan, Raees
Kim, Jaehyung
Kilic, Ozgur O.
Klasky, Scott
Klimentov, Alexei
Korchuganova, Tatiana
Outschoorn, Verena Ingrid Martinez
Nilsson, Paul
Park, David K.
Podhorszki, Norbert
Ren, Yihui
Steele, John Rembrandt
Suter, Frédéric
Vatsavai, Sairam Sri
Wenaus, Torre
Yang, Wei
Yang, Yiming
Yoo, Shinjae
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
68T05, 68M14, 68W10
The collaborative efforts of large communities in science experiments, often comprising thousands of global members, reflect a monumental commitment to exploration and discovery. Recently, advanced and complex data processing has gained increasing importance in science experiments. Data processing workflows typically consist of multiple intricate steps, and the precise specification of resource requirements is crucial for each step to allocate optimal resources for effective processing. Estimating resource requirements in advance is challenging due to a wide range of analysis scenarios, varying skill levels among community members, and the continuously increasing spectrum of computing options. One practical approach to mitigate these challenges involves initially processing a subset of each step to measure precise resource utilization from actual processing profiles before completing the entire step. While this two-staged approach enables processing on optimal resources for most of the workflow, it has drawbacks such as initial inaccuracies leading to potential failures and suboptimal resource usage, along with overhead from waiting for initial processing completion, which is critical for fast-turnaround analyses. In this context, our study introduces a novel pipeline of machine learning models within a comprehensive workflow management system, the Production and Distributed Analysis (PanDA) system. These models employ advanced machine learning techniques to predict key resource requirements, overcoming challenges posed by limited upfront knowledge of characteristics at each step. Accurate forecasts of resource requirements enable informed and proactive decision-making in workflow management, enhancing the efficiency of handling diverse, complex workflows across heterogeneous resources.
title Machine Learning-Driven Predictive Resource Management in Complex Science Workflows
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
68T05, 68M14, 68W10
url https://arxiv.org/abs/2509.11512