Scalable Runtime Architecture for Data-driven, Hybrid HPC and ML Workflow Applications
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910879077892096 |
|---|---|
| author | Merzky, Andre Titov, Mikhail Turilli, Matteo Kilic, Ozgur Wang, Tianle Jha, Shantenu |
| author_facet | Merzky, Andre Titov, Mikhail Turilli, Matteo Kilic, Ozgur Wang, Tianle Jha, Shantenu |
| contents | Hybrid workflows combining traditional HPC and novel ML methodologies are transforming scientific computing. This paper presents the architecture and implementation of a scalable runtime system that extends RADICAL-Pilot with service-based execution to support AI-out-HPC workflows. Our runtime system enables distributed ML capabilities, efficient resource management, and seamless HPC/ML coupling across local and remote platforms. Preliminary experimental results show that our approach manages concurrent execution of ML models across local and remote HPC/cloud resources with minimal architectural overheads. This lays the foundation for prototyping three representative data-driven workflow applications and executing them at scale on leadership-class HPC platforms. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_13343 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Scalable Runtime Architecture for Data-driven, Hybrid HPC and ML Workflow Applications Merzky, Andre Titov, Mikhail Turilli, Matteo Kilic, Ozgur Wang, Tianle Jha, Shantenu Distributed, Parallel, and Cluster Computing Artificial Intelligence Hybrid workflows combining traditional HPC and novel ML methodologies are transforming scientific computing. This paper presents the architecture and implementation of a scalable runtime system that extends RADICAL-Pilot with service-based execution to support AI-out-HPC workflows. Our runtime system enables distributed ML capabilities, efficient resource management, and seamless HPC/ML coupling across local and remote platforms. Preliminary experimental results show that our approach manages concurrent execution of ML models across local and remote HPC/cloud resources with minimal architectural overheads. This lays the foundation for prototyping three representative data-driven workflow applications and executing them at scale on leadership-class HPC platforms. |
| title | Scalable Runtime Architecture for Data-driven, Hybrid HPC and ML Workflow Applications |
| topic | Distributed, Parallel, and Cluster Computing Artificial Intelligence |
| url | https://arxiv.org/abs/2503.13343 |