HPAC-ML: A Programming Model for Embedding ML Surrogates in Scientific Applications
Fuente:
arXiv
Saved in:
| Main Authors: | Fink, Zane, Parasyris, Konstantinos, Rathi, Praneet, Georgakoudis, Giorgis, Menon, Harshitha, Bremer, Peer-Timo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Taking GPU Programming Models to Task for Performance Portability
by: Davis, Joshua H., et al.
Published: (2024)
by: Davis, Joshua H., et al.
Published: (2024)
Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search
by: Nichols, Daniel, et al.
Published: (2026)
by: Nichols, Daniel, et al.
Published: (2026)
Counting Without Running: Evaluating LLMs' Reasoning About Code Complexity
by: Bolet, Gregory, et al.
Published: (2025)
by: Bolet, Gregory, et al.
Published: (2025)
Can Large Language Models Predict Parallel Code Performance?
by: Bolet, Gregory, et al.
Published: (2025)
by: Bolet, Gregory, et al.
Published: (2025)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
by: Nichols, Daniel, et al.
Published: (2025)
by: Nichols, Daniel, et al.
Published: (2025)
LLMs as Packagers of HPC Software
by: Melone, Caetano, et al.
Published: (2025)
by: Melone, Caetano, et al.
Published: (2025)
HPC-Coder: Modeling Parallel Programs using Large Language Models
by: Nichols, Daniel, et al.
Published: (2023)
by: Nichols, Daniel, et al.
Published: (2023)
Declarative Data Pipeline for Large Scale ML Services
by: Yang, Yunzhao, et al.
Published: (2025)
by: Yang, Yunzhao, et al.
Published: (2025)
ML-based Modeling to Predict I/O Performance on Different Storage Sub-systems
by: Xu, Yiheng, et al.
Published: (2023)
by: Xu, Yiheng, et al.
Published: (2023)
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
by: Svedas, Jonas, et al.
Published: (2026)
by: Svedas, Jonas, et al.
Published: (2026)
PCCL: Photonic circuit-switched collective communication for distributed ML
by: Kumar, Abhishek Vijaya, et al.
Published: (2025)
by: Kumar, Abhishek Vijaya, et al.
Published: (2025)
ML-based Adaptive Prefetching and Data Placement for US HEP Systems
by: Karanam, Venkat Sai Suman Lamba, et al.
Published: (2025)
by: Karanam, Venkat Sai Suman Lamba, et al.
Published: (2025)
A Performance Analyzer for a Public Cloud's ML-Augmented VM Allocator
by: Bostandoost, Roozbeh, et al.
Published: (2025)
by: Bostandoost, Roozbeh, et al.
Published: (2025)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
by: Mehboob, Talha, et al.
Published: (2025)
by: Mehboob, Talha, et al.
Published: (2025)
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
by: Yoo, Jinsun, et al.
Published: (2026)
by: Yoo, Jinsun, et al.
Published: (2026)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
by: Ahmad, Sohaib, et al.
Published: (2024)
by: Ahmad, Sohaib, et al.
Published: (2024)
ML-ECS: A Collaborative Multimodal Learning Framework for Edge-Cloud Synergies
by: Liu, Yuze, et al.
Published: (2026)
by: Liu, Yuze, et al.
Published: (2026)
NetSenseML: Network-Adaptive Compression for Efficient Distributed Machine Learning
by: Wang, Yisu, et al.
Published: (2025)
by: Wang, Yisu, et al.
Published: (2025)
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
by: Jain, Rutwik, et al.
Published: (2024)
by: Jain, Rutwik, et al.
Published: (2024)
CubicML: Automated ML for Large ML Systems Co-design with ML Prediction of Performance
by: Wen, Wei, et al.
Published: (2024)
by: Wen, Wei, et al.
Published: (2024)
Scalable Runtime Architecture for Data-driven, Hybrid HPC and ML Workflow Applications
by: Merzky, Andre, et al.
Published: (2025)
by: Merzky, Andre, et al.
Published: (2025)
A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
by: Jeon, Beomyeol, et al.
Published: (2024)
by: Jeon, Beomyeol, et al.
Published: (2024)
Why Atomicity Matters to AI/ML Infrastructure: Snapshots, Firmware Updates, and the Cost of the Forward-In-Time-Only Category Mistake
by: Borrill, Paul
Published: (2026)
by: Borrill, Paul
Published: (2026)
A Virtual Environment for Collaborative Inspection in Additive Manufacturing
by: Chheang, Vuthea, et al.
Published: (2024)
by: Chheang, Vuthea, et al.
Published: (2024)
Ilargi: a GPU Compatible Factorized ML Model Training Framework
by: Sun, Wenbo, et al.
Published: (2025)
by: Sun, Wenbo, et al.
Published: (2025)
AI Surrogate Model for Distributed Computing Workloads
by: Park, David K., et al.
Published: (2024)
by: Park, David K., et al.
Published: (2024)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
by: Zhao, Bohan, et al.
Published: (2025)
by: Zhao, Bohan, et al.
Published: (2025)
Autonomous Electrochemistry Platform with Real-Time Normality Testing of Voltammetry Measurements Using ML
by: Al-Najjar, Anees, et al.
Published: (2025)
by: Al-Najjar, Anees, et al.
Published: (2025)
Ariel-ML: Computing Parallelization with Embedded Rust for Neural Networks on Heterogeneous Multi-core Microcontrollers
by: Huang, Zhaolan, et al.
Published: (2025)
by: Huang, Zhaolan, et al.
Published: (2025)
IPComp: Interpolation Based Progressive Lossy Compression for Scientific Applications
by: Yang, Zhuoxun, et al.
Published: (2025)
by: Yang, Zhuoxun, et al.
Published: (2025)
Snowpark: Performant, Secure, User-Friendly Data Engineering and AI/ML Next To Your Data
by: Baker, Brandon, et al.
Published: (2025)
by: Baker, Brandon, et al.
Published: (2025)
User Experiences with MPI RMA and ULFM in a Resilient Key-Value Store Implementation
by: Fohry, Claudia, et al.
Published: (2026)
by: Fohry, Claudia, et al.
Published: (2026)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
by: Punniyamurthy, Kishore, et al.
Published: (2023)
by: Punniyamurthy, Kishore, et al.
Published: (2023)
Compass: A Decentralized Scheduler for Latency-Sensitive ML Workflows
by: Yang, Yuting, et al.
Published: (2024)
by: Yang, Yuting, et al.
Published: (2024)
Taming the Memory Beast: Strategies for Reliable ML Training on Kubernetes
by: Ray, Jaideep
Published: (2024)
by: Ray, Jaideep
Published: (2024)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
by: Agrawal, Anirudha, et al.
Published: (2024)
by: Agrawal, Anirudha, et al.
Published: (2024)
Reimagining RDMA Through the Lens of ML
by: Warraich, Ertza, et al.
Published: (2025)
by: Warraich, Ertza, et al.
Published: (2025)
Characterization of GPU TEE Overheads in Distributed Data Parallel ML Training
by: Lee, Jonghyun, et al.
Published: (2025)
by: Lee, Jonghyun, et al.
Published: (2025)
Evolving HPC services to enable ML workloads on HPE Cray EX
by: Schuppli, Stefano, et al.
Published: (2025)
by: Schuppli, Stefano, et al.
Published: (2025)
A Bring-Your-Own-Model Approach for ML-Driven Storage Placement in Warehouse-Scale Computers
by: Yang, Chenxi, et al.
Published: (2025)
by: Yang, Chenxi, et al.
Published: (2025)
Similar Items
-
Taking GPU Programming Models to Task for Performance Portability
by: Davis, Joshua H., et al.
Published: (2024) -
Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search
by: Nichols, Daniel, et al.
Published: (2026) -
Counting Without Running: Evaluating LLMs' Reasoning About Code Complexity
by: Bolet, Gregory, et al.
Published: (2025) -
Can Large Language Models Predict Parallel Code Performance?
by: Bolet, Gregory, et al.
Published: (2025) -
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
by: Nichols, Daniel, et al.
Published: (2025)