Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
Fuente:
arXiv
Saved in:
| Main Authors: | Fernandez, Jared, Wehrstedt, Luca, Shamis, Leonid, Elhoushi, Mostafa, Saladi, Kalyan, Bisk, Yonatan, Strubell, Emma, Kahn, Jacob |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
by: Yan, Ran, et al.
Published: (2024)
by: Yan, Ran, et al.
Published: (2024)
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
by: Kokolis, Apostolos, et al.
Published: (2024)
by: Kokolis, Apostolos, et al.
Published: (2024)
Optimizing Distributed Training Approaches for Scaling Neural Networks
by: Baligodugula, Vishnu Vardhan, et al.
Published: (2025)
by: Baligodugula, Vishnu Vardhan, et al.
Published: (2025)
Scheduling Data-Intensive Workloads in Large-Scale Distributed Systems: Trends and Challenges
by: Stavrinides, Georgios L., et al.
Published: (2025)
by: Stavrinides, Georgios L., et al.
Published: (2025)
The Energy Cost of Execution-Idle in GPU Clusters
by: Lei, Yiran, et al.
Published: (2026)
by: Lei, Yiran, et al.
Published: (2026)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
by: Zhang, WenZheng, et al.
Published: (2024)
by: Zhang, WenZheng, et al.
Published: (2024)
It's not a lie if you don't get caught: simplifying reconfiguration in SMR through dirty logs
by: Clement, Allen, et al.
Published: (2026)
by: Clement, Allen, et al.
Published: (2026)
RapidGNN: Communication Efficient Large-Scale Distributed Training of Graph Neural Networks
by: Niam, Arefin, et al.
Published: (2025)
by: Niam, Arefin, et al.
Published: (2025)
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
by: Golden, Alicia, et al.
Published: (2025)
by: Golden, Alicia, et al.
Published: (2025)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
by: Ahmad, Sohaib, et al.
Published: (2024)
by: Ahmad, Sohaib, et al.
Published: (2024)
Echo: Simulating Distributed Training At Scale
by: Feng, Yicheng, et al.
Published: (2024)
by: Feng, Yicheng, et al.
Published: (2024)
Scaling MPI Applications on Aurora
by: Ibeid, Huda, et al.
Published: (2025)
by: Ibeid, Huda, et al.
Published: (2025)
ACE-Sync: An Adaptive Cloud-Edge Synchronization Framework for Communication-Efficient Large-Scale Distributed Model Training
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
by: Xue, Chunyu, et al.
Published: (2026)
by: Xue, Chunyu, et al.
Published: (2026)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
by: Li, Haley, et al.
Published: (2026)
by: Li, Haley, et al.
Published: (2026)
Speedup of Distributed Algorithms for Power Graphs in the CONGEST Model
by: Barenboim, Leonid, et al.
Published: (2023)
by: Barenboim, Leonid, et al.
Published: (2023)
Efficient Parallelization Layouts for Large-Scale Distributed Model Training
by: Hagemann, Johannes, et al.
Published: (2023)
by: Hagemann, Johannes, et al.
Published: (2023)
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
by: Sun, Mo, et al.
Published: (2024)
by: Sun, Mo, et al.
Published: (2024)
Justin: Hybrid CPU/Memory Elastic Scaling for Distributed Stream Processing
by: Schmitz, Donatien, et al.
Published: (2025)
by: Schmitz, Donatien, et al.
Published: (2025)
An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience
by: Coles, Jonathan, et al.
Published: (2026)
by: Coles, Jonathan, et al.
Published: (2026)
Optimizing High-Throughput Distributed Data Pipelines for Reproducible Deep Learning at Scale
by: Mittal, Kashish, et al.
Published: (2026)
by: Mittal, Kashish, et al.
Published: (2026)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
by: Chen, Jiu, et al.
Published: (2026)
by: Chen, Jiu, et al.
Published: (2026)
EMLIO: Minimizing I/O Latency and Energy Consumption for Large-Scale AI Training
by: Jamil, Hasibul, et al.
Published: (2025)
by: Jamil, Hasibul, et al.
Published: (2025)
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning
by: Xu, Lang, et al.
Published: (2025)
by: Xu, Lang, et al.
Published: (2025)
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
by: Liang, Yan, et al.
Published: (2026)
by: Liang, Yan, et al.
Published: (2026)
Diagonal Scaling: A Multi-Dimensional Resource Model and Optimization Framework for Distributed Databases
by: Abdullah, Shahir, et al.
Published: (2025)
by: Abdullah, Shahir, et al.
Published: (2025)
madupite: A High-Performance Distributed Solver for Large-Scale Markov Decision Processes
by: Gargiani, Matilde, et al.
Published: (2025)
by: Gargiani, Matilde, et al.
Published: (2025)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
by: Adnan, Muhammad, et al.
Published: (2024)
by: Adnan, Muhammad, et al.
Published: (2024)
HETHUB: A Distributed Training System with Heterogeneous Cluster for Large-Scale Models
by: Xu, Si, et al.
Published: (2024)
by: Xu, Si, et al.
Published: (2024)
Armada: Memory-Efficient Distributed Training of Large-Scale Graph Neural Networks
by: Waleffe, Roger, et al.
Published: (2025)
by: Waleffe, Roger, et al.
Published: (2025)
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
by: Qu, Huanyu, et al.
Published: (2025)
by: Qu, Huanyu, et al.
Published: (2025)
Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices
by: Shen, Tao, et al.
Published: (2025)
by: Shen, Tao, et al.
Published: (2025)
SDSL-Solver: Scalable Distributed Sparse Linear Solvers for Large-Scale Interior Point Methods
by: Yang, Shaofeng, et al.
Published: (2026)
by: Yang, Shaofeng, et al.
Published: (2026)
StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications
by: Wen, Linfeng, et al.
Published: (2024)
by: Wen, Linfeng, et al.
Published: (2024)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
by: Yu, Minchen, et al.
Published: (2025)
by: Yu, Minchen, et al.
Published: (2025)
A Scalable Recipe on SuperMUC-NG Phase 2: Efficient Large-Scale Training of Language Models
by: Rajgopal, Ajay Navilarekal, et al.
Published: (2026)
by: Rajgopal, Ajay Navilarekal, et al.
Published: (2026)
Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
by: Li, Shengwei, et al.
Published: (2023)
by: Li, Shengwei, et al.
Published: (2023)
A Tale of Two Scales: Reconciling Horizontal and Vertical Scaling for Inference Serving Systems
by: Razavi, Kamran, et al.
Published: (2024)
by: Razavi, Kamran, et al.
Published: (2024)
MPI-Q: A Message Communication Library for Large-Scale Classical-Quantum Heterogeneous Hybrid Distributed Computing
by: Wang, Feng, et al.
Published: (2026)
by: Wang, Feng, et al.
Published: (2026)
FAIR Ecosystems for Science at Scale
by: Wilkinson, Sean R., et al.
Published: (2025)
by: Wilkinson, Sean R., et al.
Published: (2025)
Similar Items
-
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
by: Yan, Ran, et al.
Published: (2024) -
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
by: Kokolis, Apostolos, et al.
Published: (2024) -
Optimizing Distributed Training Approaches for Scaling Neural Networks
by: Baligodugula, Vishnu Vardhan, et al.
Published: (2025) -
Scheduling Data-Intensive Workloads in Large-Scale Distributed Systems: Trends and Challenges
by: Stavrinides, Georgios L., et al.
Published: (2025) -
The Energy Cost of Execution-Idle in GPU Clusters
by: Lei, Yiran, et al.
Published: (2026)