LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Mo, Yang, Zihan, Liao, Changyue, Li, Yingtao, Wu, Fei, Wang, Zeke |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LoHan: Low-Cost High-Performance Framework to Fine-Tune 100B Model on a Consumer GPU
by: Liao, Changyue, et al.
Published: (2024)
by: Liao, Changyue, et al.
Published: (2024)
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
by: Svedas, Jonas, et al.
Published: (2025)
by: Svedas, Jonas, et al.
Published: (2025)
Optimizing Distributed Training Approaches for Scaling Neural Networks
by: Baligodugula, Vishnu Vardhan, et al.
Published: (2025)
by: Baligodugula, Vishnu Vardhan, et al.
Published: (2025)
Parallelizing Large-Scale Tensor Network Contraction on Multiple GPUs
by: Pan, Feng, et al.
Published: (2026)
by: Pan, Feng, et al.
Published: (2026)
Cost-Effective Edge Data Distribution with End-To-End Delay Guarantees in Edge Computing
by: Shankar, Ravi, et al.
Published: (2025)
by: Shankar, Ravi, et al.
Published: (2025)
A Proposed End-To-End Principle for Data Commons
by: Grossman, Robert L.
Published: (2025)
by: Grossman, Robert L.
Published: (2025)
End-to-End and Phase-Level Performance Optimization for Hyperledger Fabric
by: Sollu, Pavan, et al.
Published: (2026)
by: Sollu, Pavan, et al.
Published: (2026)
ATLAS: Efficient Out-of-Core Inference for Billion-Scale Graph Neural Networks
by: Naman, Pranjal, et al.
Published: (2026)
by: Naman, Pranjal, et al.
Published: (2026)
Saarthi: An End-to-End Intelligent Platform for Optimising Distributed Serverless Workloads
by: Agarwal, Siddharth, et al.
Published: (2025)
by: Agarwal, Siddharth, et al.
Published: (2025)
Parallelizing Maximal Clique Enumeration on GPUs
by: Almasri, Mohammad, et al.
Published: (2022)
by: Almasri, Mohammad, et al.
Published: (2022)
Beyond End-to-End: Dynamic Chain Optimization for Private LLM Adaptation on the Edge
by: Wu, Yebo, et al.
Published: (2026)
by: Wu, Yebo, et al.
Published: (2026)
DeepCEE: Efficient Cross-Region Model Distributed Training System under Heterogeneous GPUs and Networks
by: Wang, Jinquan, et al.
Published: (2025)
by: Wang, Jinquan, et al.
Published: (2025)
Nezha: Breaking Multi-Rail Network Barriers for Distributed DNN Training
by: Yu, Enda, et al.
Published: (2024)
by: Yu, Enda, et al.
Published: (2024)
FpgaHub: Fpga-centric Hyper-heterogeneous Computing Platform for Big Data Analytics
by: Wang, Zeke, et al.
Published: (2025)
by: Wang, Zeke, et al.
Published: (2025)
ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency
by: Yao, Yuhang, et al.
Published: (2024)
by: Yao, Yuhang, et al.
Published: (2024)
Hyperion: Hierarchical Scheduling for Parallel LLM Acceleration in Multi-tier Networks
by: Ma, Mulei, et al.
Published: (2025)
by: Ma, Mulei, et al.
Published: (2025)
Scaling State-Space Models on Multiple GPUs with Tensor Parallelism
by: Dutt, Anurag, et al.
Published: (2026)
by: Dutt, Anurag, et al.
Published: (2026)
An Adaptive Distributed Stencil Abstraction for GPUs
by: Bhosale, Aditya, et al.
Published: (2025)
by: Bhosale, Aditya, et al.
Published: (2025)
Training LLMs with Fault Tolerant HSDP on 100,000 GPUs
by: Salpekar, Omkar, et al.
Published: (2026)
by: Salpekar, Omkar, et al.
Published: (2026)
RapidGNN: Communication Efficient Large-Scale Distributed Training of Graph Neural Networks
by: Niam, Arefin, et al.
Published: (2025)
by: Niam, Arefin, et al.
Published: (2025)
CaPGNN: Optimizing Parallel Graph Neural Network Training with Joint Caching and Resource-Aware Graph Partitioning
by: Song, Xianfeng, et al.
Published: (2025)
by: Song, Xianfeng, et al.
Published: (2025)
ClusterRCA: An End-to-End Approach for Network Fault Localization and Classification for HPC System
by: Sun, Yongqian, et al.
Published: (2025)
by: Sun, Yongqian, et al.
Published: (2025)
Deal: Distributed End-to-End GNN Inference for All Nodes
by: Chen, Shiyang, et al.
Published: (2025)
by: Chen, Shiyang, et al.
Published: (2025)
Combining Performance and Productivity: Accelerating the Network Sensing Graph Challenge with GPUs and Commodity Data Science Software
by: Samsi, Siddharth, et al.
Published: (2025)
by: Samsi, Siddharth, et al.
Published: (2025)
Optimizing sDTW for AMD GPUs
by: Latta-Lin, Daniel, et al.
Published: (2024)
by: Latta-Lin, Daniel, et al.
Published: (2024)
Anonymized Network Sensing using C++26 std::execution on GPUs
by: Mandulak, Michael, et al.
Published: (2025)
by: Mandulak, Michael, et al.
Published: (2025)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
by: Chen, Aodong, et al.
Published: (2023)
by: Chen, Aodong, et al.
Published: (2023)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
by: Fan, Jiakun, et al.
Published: (2025)
by: Fan, Jiakun, et al.
Published: (2025)
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
by: Cui, Shengkun, et al.
Published: (2025)
by: Cui, Shengkun, et al.
Published: (2025)
KubeIntellect: A Modular LLM-Orchestrated Agent Framework for End-to-End Kubernetes Management
by: Ardebili, Mohsen Seyedkazemi, et al.
Published: (2025)
by: Ardebili, Mohsen Seyedkazemi, et al.
Published: (2025)
A Parallel and Distributed Rust Library for Core Decomposition on Large Graphs
by: Rucci, Davide, et al.
Published: (2025)
by: Rucci, Davide, et al.
Published: (2025)
Heta: Distributed Training of Heterogeneous Graph Neural Networks
by: Zhong, Yuchen, et al.
Published: (2024)
by: Zhong, Yuchen, et al.
Published: (2024)
DreamDDP: Accelerating Data Parallel Distributed LLM Training with Layer-wise Scheduled Partial Synchronization
by: Tang, Zhenheng, et al.
Published: (2025)
by: Tang, Zhenheng, et al.
Published: (2025)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
by: Mo, Zizhao, et al.
Published: (2025)
by: Mo, Zizhao, et al.
Published: (2025)
SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
by: Lee, Jin, et al.
Published: (2026)
by: Lee, Jin, et al.
Published: (2026)
Parallel Collaborative ADMM Privacy Computing and Adaptive GPU Acceleration for Distributed Edge Networks
by: Xia, Mengchun, et al.
Published: (2026)
by: Xia, Mengchun, et al.
Published: (2026)
Accelerating End-Cloud Collaborative Inference via Near Bubble-free Pipeline Optimization
by: Gao, Luyao, et al.
Published: (2024)
by: Gao, Luyao, et al.
Published: (2024)
Efficient Parallelization Layouts for Large-Scale Distributed Model Training
by: Hagemann, Johannes, et al.
Published: (2023)
by: Hagemann, Johannes, et al.
Published: (2023)
Distributed Consensus Network: A Modularized Communication Framework and Reliability Probabilistic Analysis
by: Li, Yuetai, et al.
Published: (2025)
by: Li, Yuetai, et al.
Published: (2025)
Similar Items
-
LoHan: Low-Cost High-Performance Framework to Fine-Tune 100B Model on a Consumer GPU
by: Liao, Changyue, et al.
Published: (2024) -
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
by: Svedas, Jonas, et al.
Published: (2025) -
Optimizing Distributed Training Approaches for Scaling Neural Networks
by: Baligodugula, Vishnu Vardhan, et al.
Published: (2025) -
Parallelizing Large-Scale Tensor Network Contraction on Multiple GPUs
by: Pan, Feng, et al.
Published: (2026) -
Cost-Effective Edge Data Distribution with End-To-End Delay Guarantees in Edge Computing
by: Shankar, Ravi, et al.
Published: (2025)