Revisiting Reliability in Large-Scale Machine Learning Research Clusters
Fuente:
arXiv
Guardado en:
| Autores principales: | Kokolis, Apostolos, Kuchnik, Michael, Hoffman, John, Kumar, Adithya, Malani, Parth, Ma, Faye, DeVito, Zachary, Sengupta, Shubho, Saladi, Kalyan, Wu, Carole-Jean |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
por: Golden, Alicia, et al.
Publicado: (2025)
por: Golden, Alicia, et al.
Publicado: (2025)
MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems
por: Hsia, Samuel, et al.
Publicado: (2023)
por: Hsia, Samuel, et al.
Publicado: (2023)
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
por: Fernandez, Jared, et al.
Publicado: (2024)
por: Fernandez, Jared, et al.
Publicado: (2024)
Is Flash Attention Stable?
por: Golden, Alicia, et al.
Publicado: (2024)
por: Golden, Alicia, et al.
Publicado: (2024)
ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training
por: Li, Minghao, et al.
Publicado: (2026)
por: Li, Minghao, et al.
Publicado: (2026)
Revisit to the Bai-Galbraith signature scheme
por: Sengupta, Banhirup, et al.
Publicado: (2025)
por: Sengupta, Banhirup, et al.
Publicado: (2025)
Generative AI Beyond LLMs: System Implications of Multi-Modal Generation
por: Golden, Alicia, et al.
Publicado: (2023)
por: Golden, Alicia, et al.
Publicado: (2023)
Automated, Reliable, and Efficient Continental-Scale Replication of 7.3 Petabytes of Climate Simulation Data: A Case Study
por: Lacinski, Lukasz, et al.
Publicado: (2024)
por: Lacinski, Lukasz, et al.
Publicado: (2024)
An Overview on the Landscape of Self-Adaptive Cloud Design and Operation Patterns: Goals, Strategies, Tooling, Evaluation, and Dataset Perspectives
por: Angelis, Apostolos, et al.
Publicado: (2025)
por: Angelis, Apostolos, et al.
Publicado: (2025)
Large Scale Multi-GPU Based Parallel Traffic Simulation for Accelerated Traffic Assignment and Propagation
por: Jiang, Xuan, et al.
Publicado: (2024)
por: Jiang, Xuan, et al.
Publicado: (2024)
Beyond Efficiency: Scaling AI Sustainably
por: Wu, Carole-Jean, et al.
Publicado: (2024)
por: Wu, Carole-Jean, et al.
Publicado: (2024)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
por: Zhang, Mingjun, et al.
Publicado: (2025)
por: Zhang, Mingjun, et al.
Publicado: (2025)
AIReSim: A Discrete Event Simulator for Large-scale AI Cluster Reliability Modeling
por: Pattabiraman, Karthik, et al.
Publicado: (2026)
por: Pattabiraman, Karthik, et al.
Publicado: (2026)
Scaling MPI Applications on Aurora
por: Ibeid, Huda, et al.
Publicado: (2025)
por: Ibeid, Huda, et al.
Publicado: (2025)
Carbon: Scaling Trusted Payments with Untrusted Machines
por: Camaioni, Martina, et al.
Publicado: (2022)
por: Camaioni, Martina, et al.
Publicado: (2022)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
por: Zhang, WenZheng, et al.
Publicado: (2024)
por: Zhang, WenZheng, et al.
Publicado: (2024)
PRISM: Dynamic Primitive-Based Forecasting for Large-Scale GPU Cluster Workloads
por: Wu, Xin, et al.
Publicado: (2026)
por: Wu, Xin, et al.
Publicado: (2026)
Understanding Large-Scale HPC System Behavior Through Cluster-Based Visual Analytics
por: Austin, Allison, et al.
Publicado: (2026)
por: Austin, Allison, et al.
Publicado: (2026)
M$^2$-MFP: A Multi-Scale and Multi-Level Memory Failure Prediction Framework for Reliable Cloud Infrastructure
por: Xie, Hongyi, et al.
Publicado: (2025)
por: Xie, Hongyi, et al.
Publicado: (2025)
Big Data Architecture for Large Organizations
por: Ismail, Fathima Nuzla, et al.
Publicado: (2025)
por: Ismail, Fathima Nuzla, et al.
Publicado: (2025)
Mao: Machine learning approach for NUMA optimization in Warehouse Scale Computers
por: Liu, Yueji, et al.
Publicado: (2024)
por: Liu, Yueji, et al.
Publicado: (2024)
Ira: Efficient Transaction Replay for Distributed Systems
por: Bhat, Adithya, et al.
Publicado: (2026)
por: Bhat, Adithya, et al.
Publicado: (2026)
Dynamic Probabilistic Reliable Broadcast
por: Anikina, Veronika, et al.
Publicado: (2023)
por: Anikina, Veronika, et al.
Publicado: (2023)
Scaling Up Throughput-oriented LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
por: Phung, Thanh Son, et al.
Publicado: (2025)
por: Phung, Thanh Son, et al.
Publicado: (2025)
Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo
por: Charles, Zachary, et al.
Publicado: (2025)
por: Charles, Zachary, et al.
Publicado: (2025)
Polynomial Time Local Decision Revisited
por: Feuilloley, Laurent, et al.
Publicado: (2026)
por: Feuilloley, Laurent, et al.
Publicado: (2026)
Sustaining Exascale Performance: Lessons from HPL and HPL-MxP on Aurora
por: Goto, Kazushige, et al.
Publicado: (2026)
por: Goto, Kazushige, et al.
Publicado: (2026)
H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips
por: Tang, Ding, et al.
Publicado: (2025)
por: Tang, Ding, et al.
Publicado: (2025)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
por: Wang, Yuxin, et al.
Publicado: (2023)
por: Wang, Yuxin, et al.
Publicado: (2023)
Reliable Replication Protocols on SmartNICs
por: Katebzadeh, M. R. Siavash, et al.
Publicado: (2025)
por: Katebzadeh, M. R. Siavash, et al.
Publicado: (2025)
Revisiting the Time Cost Model of AllReduce
por: Xiong, Dian, et al.
Publicado: (2024)
por: Xiong, Dian, et al.
Publicado: (2024)
Revisiting Lower Bounds for Two-Step Consensus
por: Ryabinin, Fedor, et al.
Publicado: (2025)
por: Ryabinin, Fedor, et al.
Publicado: (2025)
Reliable Communication in Hybrid Authentication and Trust Models
por: Chotkan, Rowdy, et al.
Publicado: (2024)
por: Chotkan, Rowdy, et al.
Publicado: (2024)
Introduction to Number Theoretic Transform
por: Sengupta, Banhirup, et al.
Publicado: (2025)
por: Sengupta, Banhirup, et al.
Publicado: (2025)
Applying Large-Scale Distributed Computing to Structural Bioinformatics -- Bridging Legacy HPC Clusters With Big Data Technologies Using kafka-slurm-agent
por: Rubach, Pawel
Publicado: (2025)
por: Rubach, Pawel
Publicado: (2025)
On the Solvability of Byzantine-tolerant Reliable Communication in Dynamic Networks
por: Bonomi, Silvia, et al.
Publicado: (2025)
por: Bonomi, Silvia, et al.
Publicado: (2025)
Sparse Checkpointing for Fast and Reliable MoE Training
por: Gandhi, Swapnil, et al.
Publicado: (2024)
por: Gandhi, Swapnil, et al.
Publicado: (2024)
Amortized Asynchronous Byzantine Reliable Broadcast with Optimal Resilience
por: Hu, Michael Yiqing, et al.
Publicado: (2026)
por: Hu, Michael Yiqing, et al.
Publicado: (2026)
Optimistic, Signature-Free Reliable Broadcast and Its Applications
por: Shrestha, Nibesh, et al.
Publicado: (2025)
por: Shrestha, Nibesh, et al.
Publicado: (2025)
Byzantine Reliable Broadcast with Low Communication and Time Complexity
por: Locher, Thomas
Publicado: (2024)
por: Locher, Thomas
Publicado: (2024)
Ejemplares similares
-
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
por: Golden, Alicia, et al.
Publicado: (2025) -
MAD Max Beyond Single-Node: Enabling Large Machine Learning Model Acceleration on Distributed Systems
por: Hsia, Samuel, et al.
Publicado: (2023) -
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
por: Fernandez, Jared, et al.
Publicado: (2024) -
Is Flash Attention Stable?
por: Golden, Alicia, et al.
Publicado: (2024) -
ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training
por: Li, Minghao, et al.
Publicado: (2026)