Steering a Fleet: Adaptation for Large-Scale, Workflow-Based Experiments
Fuente:
arXiv
Saved in:
| Main Authors: | Pruyne, Jim, Hayot-Sasson, Valerie, Zheng, Weijian, Chard, Ryan, Wozniak, Justin M., Bicer, Tekin, Chard, Kyle, Foster, Ian T. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Employing Artificial Intelligence to Steer Exascale Workflows with Colmena
by: Ward, Logan, et al.
Published: (2024)
by: Ward, Logan, et al.
Published: (2024)
Accelerating Python Applications with Dask and ProxyStore
by: Pauloski, J. Gregory, et al.
Published: (2024)
by: Pauloski, J. Gregory, et al.
Published: (2024)
WRATH: Workload Resilience Across Task Hierarchies in Task-based Parallel Programming Frameworks
by: Zhou, Sicheng, et al.
Published: (2025)
by: Zhou, Sicheng, et al.
Published: (2025)
Flight: A FaaS-Based Framework for Complex and Hierarchical Federated Learning
by: Hudson, Nathaniel, et al.
Published: (2024)
by: Hudson, Nathaniel, et al.
Published: (2024)
Octopus: Experiences with a Hybrid Event-Driven Architecture for Distributed Scientific Computing
by: Pan, Haochen, et al.
Published: (2024)
by: Pan, Haochen, et al.
Published: (2024)
Addressing Reproducibility Challenges in HPC with Continuous Integration
by: Hayot-Sasson, Valérie, et al.
Published: (2025)
by: Hayot-Sasson, Valérie, et al.
Published: (2025)
Object Proxy Patterns for Accelerating Distributed Applications
by: Pauloski, J. Gregory, et al.
Published: (2024)
by: Pauloski, J. Gregory, et al.
Published: (2024)
TaPS: A Performance Evaluation Suite for Task-based Execution Frameworks
by: Pauloski, J. Gregory, et al.
Published: (2024)
by: Pauloski, J. Gregory, et al.
Published: (2024)
UniFaaS: Programming across Distributed Cyberinfrastructure with Federated Function Serving
by: Li, Yifei, et al.
Published: (2024)
by: Li, Yifei, et al.
Published: (2024)
XaaS Containers: Performance-Portable Representation With Source and IR Containers
by: Copik, Marcin, et al.
Published: (2025)
by: Copik, Marcin, et al.
Published: (2025)
GreenFaaS: Maximizing Energy Efficiency of HPC Workloads with FaaS
by: Kamatar, Alok, et al.
Published: (2024)
by: Kamatar, Alok, et al.
Published: (2024)
Core Hours and Carbon Credits: Incentivizing Sustainability in HPC
by: Kamatar, Alok, et al.
Published: (2025)
by: Kamatar, Alok, et al.
Published: (2025)
Icicle: Scalable Metadata Indexing and Real-Time Monitoring for HPC File Systems
by: Pan, Haochen, et al.
Published: (2026)
by: Pan, Haochen, et al.
Published: (2026)
Empowering Scientific Workflows with Federated Agents
by: Kamatar, Alok, et al.
Published: (2025)
by: Kamatar, Alok, et al.
Published: (2025)
DynoStore: A wide-area distribution system for the management of data over heterogeneous storage
by: Sanchez-Gallegos, Dante D., et al.
Published: (2025)
by: Sanchez-Gallegos, Dante D., et al.
Published: (2025)
Hierarchical storage management in user space for neuroimaging applications
by: Hayot-Sasson, Valérie, et al.
Published: (2024)
by: Hayot-Sasson, Valérie, et al.
Published: (2024)
Parsl+CWL: Towards Combining the Python and CWL Ecosystems
by: Karle, Nishchay, et al.
Published: (2024)
by: Karle, Nishchay, et al.
Published: (2024)
D-Rex: Heterogeneity-Aware Reliability Framework and Adaptive Algorithms for Distributed Storage
by: Gonthier, Maxime, et al.
Published: (2025)
by: Gonthier, Maxime, et al.
Published: (2025)
Performance comparison of Dask and Apache Spark on HPC systems for Neuroimaging
by: Dugré, Mathieu, et al.
Published: (2024)
by: Dugré, Mathieu, et al.
Published: (2024)
Topology-Aware Knowledge Propagation in Decentralized Learning
by: Sakarvadia, Mansi, et al.
Published: (2025)
by: Sakarvadia, Mansi, et al.
Published: (2025)
Experiences with Model Context Protocol Servers for Science and High Performance Computing
by: Pan, Haochen, et al.
Published: (2025)
by: Pan, Haochen, et al.
Published: (2025)
Automated, Reliable, and Efficient Continental-Scale Replication of 7.3 Petabytes of Climate Simulation Data: A Case Study
by: Lacinski, Lukasz, et al.
Published: (2024)
by: Lacinski, Lukasz, et al.
Published: (2024)
PSI/J: A Portable Interface for Submitting, Monitoring, and Managing Jobs
by: Hategan-Marandiuc, Mihael, et al.
Published: (2023)
by: Hategan-Marandiuc, Mihael, et al.
Published: (2023)
Experiences Building Enterprise-Level Privacy-Preserving Federated Learning to Power AI for Science
by: Li, Zilinghan, et al.
Published: (2025)
by: Li, Zilinghan, et al.
Published: (2025)
Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
by: Wang, Wenyi, et al.
Published: (2025)
by: Wang, Wenyi, et al.
Published: (2025)
A Terminology for Scientific Workflow Systems
by: Suter, Frédéric, et al.
Published: (2025)
by: Suter, Frédéric, et al.
Published: (2025)
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
by: Ma, Bin, et al.
Published: (2026)
by: Ma, Bin, et al.
Published: (2026)
ExaWorks Software Development Kit: A Robust and Scalable Collection of Interoperable Workflow Technologies
by: Turilli, Matteo, et al.
Published: (2024)
by: Turilli, Matteo, et al.
Published: (2024)
mLR: Scalable Laminography Reconstruction based on Memoization
by: Ma, Bin, et al.
Published: (2025)
by: Ma, Bin, et al.
Published: (2025)
Exploring Distributed Vector Databases Performance on HPC Platforms: A Study with Qdrant
by: Ockerman, Seth, et al.
Published: (2025)
by: Ockerman, Seth, et al.
Published: (2025)
Trillion Parameter AI Serving Infrastructure for Scientific Discovery: A Survey and Vision
by: Hudson, Nathaniel, et al.
Published: (2024)
by: Hudson, Nathaniel, et al.
Published: (2024)
MOFA: Discovering Materials for Carbon Capture with a GenAI- and Simulation-Based Workflow
by: Yan, Xiaoli, et al.
Published: (2025)
by: Yan, Xiaoli, et al.
Published: (2025)
FleetOpt: Analytical Fleet Provisioning for LLM Inference with Compress-and-Route as Implementation Mechanism
by: Chen, Huamin, et al.
Published: (2026)
by: Chen, Huamin, et al.
Published: (2026)
Computational Grids
by: Foster, Ian, et al.
Published: (2025)
by: Foster, Ian, et al.
Published: (2025)
Workflows Community Summit 2024: Future Trends and Challenges in Scientific Workflows
by: da Silva, Rafael Ferreira, et al.
Published: (2024)
by: da Silva, Rafael Ferreira, et al.
Published: (2024)
Enabling Seamless Transitions from Experimental to Production HPC for Interactive Workflows
by: Etz, Brian D., et al.
Published: (2025)
by: Etz, Brian D., et al.
Published: (2025)
Towards Experiment Execution in Support of Community Benchmark Workflows for HPC
by: von Laszewski, Gregor, et al.
Published: (2025)
by: von Laszewski, Gregor, et al.
Published: (2025)
RHAPSODY: Execution of Hybrid AI-HPC Workflows at Scale
by: Alsaadi, Aymen, et al.
Published: (2025)
by: Alsaadi, Aymen, et al.
Published: (2025)
FIRST: Federated Inference Resource Scheduling Toolkit for Scientific AI Model Access
by: Tanikanti, Aditya, et al.
Published: (2025)
by: Tanikanti, Aditya, et al.
Published: (2025)
It Takes Two to Tango: Serverless Workflow Serving via Bilaterally Engaged Resource Adaptation
by: Wu, Jing, et al.
Published: (2025)
by: Wu, Jing, et al.
Published: (2025)
Similar Items
-
Employing Artificial Intelligence to Steer Exascale Workflows with Colmena
by: Ward, Logan, et al.
Published: (2024) -
Accelerating Python Applications with Dask and ProxyStore
by: Pauloski, J. Gregory, et al.
Published: (2024) -
WRATH: Workload Resilience Across Task Hierarchies in Task-based Parallel Programming Frameworks
by: Zhou, Sicheng, et al.
Published: (2025) -
Flight: A FaaS-Based Framework for Complex and Hierarchical Federated Learning
by: Hudson, Nathaniel, et al.
Published: (2024) -
Octopus: Experiences with a Hybrid Event-Driven Architecture for Distributed Scientific Computing
by: Pan, Haochen, et al.
Published: (2024)