A Unifying Framework to Enable Artificial Intelligence in High Performance Computing Workflows
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908350622466048 |
|---|---|
| author | Domke, Jens Wahib, Mohamed Dubey, Anshu Ben-Nun, Tal Draeger, Erik W. |
| author_facet | Domke, Jens Wahib, Mohamed Dubey, Anshu Ben-Nun, Tal Draeger, Erik W. |
| contents | Current trends point to a future where large-scale scientific applications are tightly-coupled HPC/AI hybrids. Hence, we urgently need to invest in creating a seamless, scalable framework where HPC and AI/ML can efficiently work together and adapt to novel hardware and vendor libraries without starting from scratch every few years. The current ecosystem and sparsely-connected community are not sufficient to tackle these challenges, and we require a breakthrough catalyst for science similar to what PyTorch enabled for AI. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_02738 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Unifying Framework to Enable Artificial Intelligence in High Performance Computing Workflows Domke, Jens Wahib, Mohamed Dubey, Anshu Ben-Nun, Tal Draeger, Erik W. Distributed, Parallel, and Cluster Computing Software Engineering Current trends point to a future where large-scale scientific applications are tightly-coupled HPC/AI hybrids. Hence, we urgently need to invest in creating a seamless, scalable framework where HPC and AI/ML can efficiently work together and adapt to novel hardware and vendor libraries without starting from scratch every few years. The current ecosystem and sparsely-connected community are not sufficient to tackle these challenges, and we require a breakthrough catalyst for science similar to what PyTorch enabled for AI. |
| title | A Unifying Framework to Enable Artificial Intelligence in High Performance Computing Workflows |
| topic | Distributed, Parallel, and Cluster Computing Software Engineering |
| url | https://arxiv.org/abs/2505.02738 |