Alps, a versatile research infrastructure
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Martinasso, Maxime, Klein, Mark, Schulthess, Thomas C. |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience
par: Coles, Jonathan, et autres
Publié: (2026)
par: Coles, Jonathan, et autres
Publié: (2026)
Beyond Pre-Training: The Full Lifecycle of Foundation Models on HPC Systems
par: Conciatore, Dino, et autres
Publié: (2026)
par: Conciatore, Dino, et autres
Publié: (2026)
Evolving HPC services to enable ML workloads on HPE Cray EX
par: Schuppli, Stefano, et autres
Publié: (2025)
par: Schuppli, Stefano, et autres
Publié: (2025)
Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
par: Fusco, Luigi, et autres
Publié: (2024)
par: Fusco, Luigi, et autres
Publié: (2024)
Simulating LLM training workloads for heterogeneous compute and network infrastructure
par: Kumar, Sumit, et autres
Publié: (2025)
par: Kumar, Sumit, et autres
Publié: (2025)
XaaS: Acceleration as a Service to Enable Productive High-Performance Cloud Computing
par: Hoefler, Torsten, et autres
Publié: (2024)
par: Hoefler, Torsten, et autres
Publié: (2024)
Interactive and Urgent HPC: State of the Research
par: Reuther, Albert, et autres
Publié: (2026)
par: Reuther, Albert, et autres
Publié: (2026)
Interactive and Urgent HPC: Challenges and Opportunities
par: Reuther, Albert, et autres
Publié: (2024)
par: Reuther, Albert, et autres
Publié: (2024)
Quantifying Liveness and Safety of Avalanche's Snowball
par: Kniep, Quentin, et autres
Publié: (2024)
par: Kniep, Quentin, et autres
Publié: (2024)
Workload Buoyancy: Keeping Apps Afloat by Identifying Shared Resource Bottlenecks
par: Larsson, Oliver, et autres
Publié: (2026)
par: Larsson, Oliver, et autres
Publié: (2026)
Failures of public key infrastructure: 53 year survey
par: Dumitrescu, Adrian-Tudor, et autres
Publié: (2024)
par: Dumitrescu, Adrian-Tudor, et autres
Publié: (2024)
The infrastructure powering IBM's Gen AI model development
par: Gershon, Talia, et autres
Publié: (2024)
par: Gershon, Talia, et autres
Publié: (2024)
Malicious node aware wireless multi hop networks: a systematic review of the literature and recommendations for future research
par: Pourdehghan, Shahram, et autres
Publié: (2025)
par: Pourdehghan, Shahram, et autres
Publié: (2025)
Improved Methods of Task Assignment and Resource Allocation with Preemption in Edge Computing Systems
par: Rublein, Caroline, et autres
Publié: (2024)
par: Rublein, Caroline, et autres
Publié: (2024)
bittide: Control Time, Not Flows
par: Bastiaan, Martijn, et autres
Publié: (2025)
par: Bastiaan, Martijn, et autres
Publié: (2025)
Octopus: Experiences with a Hybrid Event-Driven Architecture for Distributed Scientific Computing
par: Pan, Haochen, et autres
Publié: (2024)
par: Pan, Haochen, et autres
Publié: (2024)
Performance analysis of mdx II: A next-generation cloud platform for cross-disciplinary data science research
par: Takahashi, Keichi, et autres
Publié: (2025)
par: Takahashi, Keichi, et autres
Publié: (2025)
Embedded Made Easy -- Rethinking Embedded + Cloud Software Development (WIP)
par: Arnold, Anthony, et autres
Publié: (2026)
par: Arnold, Anthony, et autres
Publié: (2026)
A Privacy-Preserving Ecosystem for Developing Machine Learning Algorithms Using Patient Data: Insights from the TUM.ai Makeathon
par: Süwer, Simon, et autres
Publié: (2025)
par: Süwer, Simon, et autres
Publié: (2025)
TaPS: A Performance Evaluation Suite for Task-based Execution Frameworks
par: Pauloski, J. Gregory, et autres
Publié: (2024)
par: Pauloski, J. Gregory, et autres
Publié: (2024)
Advancing RT Core-Accelerated Fixed-Radius Nearest Neighbor Search
par: Meneses, Enzo, et autres
Publié: (2026)
par: Meneses, Enzo, et autres
Publié: (2026)
The workflow motif: a widely-useful performance diagnosis abstraction for distributed applications
par: Abdi, Mania, et autres
Publié: (2025)
par: Abdi, Mania, et autres
Publié: (2025)
Core Hours and Carbon Credits: Incentivizing Sustainability in HPC
par: Kamatar, Alok, et autres
Publié: (2025)
par: Kamatar, Alok, et autres
Publié: (2025)
WRATH: Workload Resilience Across Task Hierarchies in Task-based Parallel Programming Frameworks
par: Zhou, Sicheng, et autres
Publié: (2025)
par: Zhou, Sicheng, et autres
Publié: (2025)
LOCO: Rethinking Objects for Network Memory
par: Hodgkins, George, et autres
Publié: (2025)
par: Hodgkins, George, et autres
Publié: (2025)
The First OpenFOAM HPC Challenge (OHC-1)
par: Lesnik, Sergey, et autres
Publié: (2026)
par: Lesnik, Sergey, et autres
Publié: (2026)
Byzantine Reliable Broadcast with Low Communication and Time Complexity
par: Locher, Thomas
Publié: (2024)
par: Locher, Thomas
Publié: (2024)
Clock2Q+: A Simple and Efficient Replacement Algorithm for Metadata Cache in VMware vSAN
par: Zhai, Yiyan, et autres
Publié: (2025)
par: Zhai, Yiyan, et autres
Publié: (2025)
DynoStore: A wide-area distribution system for the management of data over heterogeneous storage
par: Sanchez-Gallegos, Dante D., et autres
Publié: (2025)
par: Sanchez-Gallegos, Dante D., et autres
Publié: (2025)
Artifact for A Non-Intrusive Framework for Deferred Integration of Cloud Patterns in Energy-Efficient Data-Sharing Pipelines
par: Masoudi, Sepideh, et autres
Publié: (2025)
par: Masoudi, Sepideh, et autres
Publié: (2025)
Enhancing Energy Efficiency in Scientific Workflows through CFD based PIVAEs
par: Zahir, Ali, et autres
Publié: (2026)
par: Zahir, Ali, et autres
Publié: (2026)
D-Rex: Heterogeneity-Aware Reliability Framework and Adaptive Algorithms for Distributed Storage
par: Gonthier, Maxime, et autres
Publié: (2025)
par: Gonthier, Maxime, et autres
Publié: (2025)
Lumos: Performance Characterization of WebAssembly as a Serverless Runtime in the Edge-Cloud Continuum
par: Marcelino, Cynthia, et autres
Publié: (2025)
par: Marcelino, Cynthia, et autres
Publié: (2025)
A Non-Intrusive Framework for Deferred Integration of Cloud Patterns in Energy-Efficient Data-Sharing Pipelines
par: Masoudi, Sepideh, et autres
Publié: (2025)
par: Masoudi, Sepideh, et autres
Publié: (2025)
Asymptotic Subspace Consensus in Dynamic Networks
par: Függer, Matthias, et autres
Publié: (2026)
par: Függer, Matthias, et autres
Publié: (2026)
Experiences with Model Context Protocol Servers for Science and High Performance Computing
par: Pan, Haochen, et autres
Publié: (2025)
par: Pan, Haochen, et autres
Publié: (2025)
Evaluating Versal AI Engines for option price discovery in market risk analysis
par: Klaisoongnoen, Mark, et autres
Publié: (2024)
par: Klaisoongnoen, Mark, et autres
Publié: (2024)
Roadrunner: Accelerating Data Delivery to WebAssembly-Based Serverless Functions
par: Marcelino, Cynthia, et autres
Publié: (2025)
par: Marcelino, Cynthia, et autres
Publié: (2025)
Topological Characterization of Consensus in Distributed Systems
par: Nowak, Thomas, et autres
Publié: (2019)
par: Nowak, Thomas, et autres
Publié: (2019)
HyperDrive: Scheduling Serverless Functions in the Edge-Cloud-Space 3D Continuum
par: Pusztai, Thomas, et autres
Publié: (2024)
par: Pusztai, Thomas, et autres
Publié: (2024)
Documents similaires
-
An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience
par: Coles, Jonathan, et autres
Publié: (2026) -
Beyond Pre-Training: The Full Lifecycle of Foundation Models on HPC Systems
par: Conciatore, Dino, et autres
Publié: (2026) -
Evolving HPC services to enable ML workloads on HPE Cray EX
par: Schuppli, Stefano, et autres
Publié: (2025) -
Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
par: Fusco, Luigi, et autres
Publié: (2024) -
Simulating LLM training workloads for heterogeneous compute and network infrastructure
par: Kumar, Sumit, et autres
Publié: (2025)