Scalable Runtime Architecture for Data-driven, Hybrid HPC and ML Workflow Applications

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Merzky, Andre, Titov, Mikhail, Turilli, Matteo, Kilic, Ozgur, Wang, Tianle, Jha, Shantenu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910879077892096
author Merzky, Andre
Titov, Mikhail
Turilli, Matteo
Kilic, Ozgur
Wang, Tianle
Jha, Shantenu
author_facet Merzky, Andre
Titov, Mikhail
Turilli, Matteo
Kilic, Ozgur
Wang, Tianle
Jha, Shantenu
contents Hybrid workflows combining traditional HPC and novel ML methodologies are transforming scientific computing. This paper presents the architecture and implementation of a scalable runtime system that extends RADICAL-Pilot with service-based execution to support AI-out-HPC workflows. Our runtime system enables distributed ML capabilities, efficient resource management, and seamless HPC/ML coupling across local and remote platforms. Preliminary experimental results show that our approach manages concurrent execution of ML models across local and remote HPC/cloud resources with minimal architectural overheads. This lays the foundation for prototyping three representative data-driven workflow applications and executing them at scale on leadership-class HPC platforms.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13343
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scalable Runtime Architecture for Data-driven, Hybrid HPC and ML Workflow Applications
Merzky, Andre
Titov, Mikhail
Turilli, Matteo
Kilic, Ozgur
Wang, Tianle
Jha, Shantenu
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Hybrid workflows combining traditional HPC and novel ML methodologies are transforming scientific computing. This paper presents the architecture and implementation of a scalable runtime system that extends RADICAL-Pilot with service-based execution to support AI-out-HPC workflows. Our runtime system enables distributed ML capabilities, efficient resource management, and seamless HPC/ML coupling across local and remote platforms. Preliminary experimental results show that our approach manages concurrent execution of ML models across local and remote HPC/cloud resources with minimal architectural overheads. This lays the foundation for prototyping three representative data-driven workflow applications and executing them at scale on leadership-class HPC platforms.
title Scalable Runtime Architecture for Data-driven, Hybrid HPC and ML Workflow Applications
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2503.13343