Lion Cub: Minimizing Communication Overhead in Distributed Lion
Fuente:
arXiv
Saved in:
| Main Authors: | Ishikawa, Satoki, Ben-Nun, Tal, Van Essen, Brian, Yokota, Rio, Dryden, Nikoli |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Communication Efficient Distributed Training with Distributed Lion
by: Liu, Bo, et al.
Published: (2024)
by: Liu, Bo, et al.
Published: (2024)
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
by: Li, Shigang, et al.
Published: (2020)
by: Li, Shigang, et al.
Published: (2020)
Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
by: Li, Shigang, et al.
Published: (2019)
by: Li, Shigang, et al.
Published: (2019)
Lion: Minimizing Distributed Transactions through Adaptive Replica Provision (Extended Version)
by: Zheng, Qiushi, et al.
Published: (2024)
by: Zheng, Qiushi, et al.
Published: (2024)
Accelerating Large Language Model Training with 4D Parallelism and Memory Consumption Estimator
by: Fujii, Kazuki, et al.
Published: (2024)
by: Fujii, Kazuki, et al.
Published: (2024)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
by: Zhuang, Chen, et al.
Published: (2024)
by: Zhuang, Chen, et al.
Published: (2024)
Reducing Communication Overhead in Federated Learning for Network Anomaly Detection with Adaptive Client Selection
by: Marfo, William, et al.
Published: (2025)
by: Marfo, William, et al.
Published: (2025)
FedPM: Federated Learning Using Second-order Optimization with Preconditioned Mixing of Local Parameters
by: Ishii, Hiro, et al.
Published: (2025)
by: Ishii, Hiro, et al.
Published: (2025)
Record-Remix-Replay: Hierarchical GPU Kernel Optimization using Evolutionary Search
by: Nichols, Daniel, et al.
Published: (2026)
by: Nichols, Daniel, et al.
Published: (2026)
Inductive Loop Analysis for Practical HPC Application Optimization
by: Schaad, Philipp, et al.
Published: (2025)
by: Schaad, Philipp, et al.
Published: (2025)
Near-Zero-Overhead Freshness for Recommendation Systems via Inference-Side Model Updates
by: Yu, Wenjun, et al.
Published: (2025)
by: Yu, Wenjun, et al.
Published: (2025)
An inherently parallel H2-ULV factorization for solving dense linear systems on GPUs
by: Ma, Qianxiang, et al.
Published: (2025)
by: Ma, Qianxiang, et al.
Published: (2025)
QSync: Quantization-Minimized Synchronous Distributed Training Across Hybrid Devices
by: Zhao, Juntao, et al.
Published: (2024)
by: Zhao, Juntao, et al.
Published: (2024)
SpaFL: Communication-Efficient Federated Learning with Sparse Models and Low computational Overhead
by: Kim, Minsu, et al.
Published: (2024)
by: Kim, Minsu, et al.
Published: (2024)
Communication-Efficient Federated Group Distributionally Robust Optimization
by: Guo, Zhishuai, et al.
Published: (2024)
by: Guo, Zhishuai, et al.
Published: (2024)
Communication Optimization for Distributed Training: Architecture, Advances, and Opportunities
by: Wei, Yunze, et al.
Published: (2024)
by: Wei, Yunze, et al.
Published: (2024)
SignMuon: Communication-Efficient Distributed Muon Optimization
by: Mishra, Neel, et al.
Published: (2026)
by: Mishra, Neel, et al.
Published: (2026)
The Sherpa.ai Blind Vertical Federated Learning Paradigm to Minimize the Number of Communications
by: Acero, Alex, et al.
Published: (2025)
by: Acero, Alex, et al.
Published: (2025)
Communication-Efficient Distributed Learning with Local Immediate Error Compensation
by: Cheng, Yifei, et al.
Published: (2024)
by: Cheng, Yifei, et al.
Published: (2024)
TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition
by: Vellaisamy, Prabhu, et al.
Published: (2026)
by: Vellaisamy, Prabhu, et al.
Published: (2026)
Communication-Efficient Distributed Deep Learning via Federated Dynamic Averaging
by: Theologitis, Michail, et al.
Published: (2024)
by: Theologitis, Michail, et al.
Published: (2024)
Detection of Global Anomalies on Distributed IoT Edges with Device-to-Device Communication
by: Ochiai, Hideya, et al.
Published: (2024)
by: Ochiai, Hideya, et al.
Published: (2024)
Low-Communication Resilient Distributed Estimation Algorithm Based on Memory Mechanism
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
by: Gond, Raja, et al.
Published: (2025)
by: Gond, Raja, et al.
Published: (2025)
SparDL: Distributed Deep Learning Training with Efficient Sparse Communication
by: Zhao, Minjun, et al.
Published: (2023)
by: Zhao, Minjun, et al.
Published: (2023)
Byzantine-Robust and Communication-Efficient Distributed Learning via Compressed Momentum Filtering
by: Liu, Changxin, et al.
Published: (2024)
by: Liu, Changxin, et al.
Published: (2024)
High-Dimensional Distributed Sparse Classification with Scalable Communication-Efficient Global Updates
by: Lu, Fred, et al.
Published: (2024)
by: Lu, Fred, et al.
Published: (2024)
An All-Reduce Compatible Top-K Compressor for Communication-Efficient Distributed Learning
by: Chen, Chuyan, et al.
Published: (2025)
by: Chen, Chuyan, et al.
Published: (2025)
Communication-Efficient Distributed Training for Collaborative Flat Optima Recovery in Deep Learning
by: Dimlioglu, Tolga, et al.
Published: (2025)
by: Dimlioglu, Tolga, et al.
Published: (2025)
DGEMM on Integer Matrix Multiplication Unit
by: Ootomo, Hiroyuki, et al.
Published: (2023)
by: Ootomo, Hiroyuki, et al.
Published: (2023)
OpenDiLoCo: An Open-Source Framework for Globally Distributed Low-Communication Training
by: Jaghouar, Sami, et al.
Published: (2024)
by: Jaghouar, Sami, et al.
Published: (2024)
ChunkFlow: Communication-Aware Chunked Prefetching for Layerwise Offloading in Distributed Diffusion Transformer Inference
by: Meng, Han, et al.
Published: (2026)
by: Meng, Han, et al.
Published: (2026)
Diagnosing Overhead in Dispatch Operations: Cross-architecture Observatory
by: Ma, Bole, et al.
Published: (2026)
by: Ma, Bole, et al.
Published: (2026)
A Unifying Framework to Enable Artificial Intelligence in High Performance Computing Workflows
by: Domke, Jens, et al.
Published: (2025)
by: Domke, Jens, et al.
Published: (2025)
Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers
by: Maczan, Jędrzej
Published: (2026)
by: Maczan, Jędrzej
Published: (2026)
CDFGNN: a Systematic Design of Cache-based Distributed Full-Batch Graph Neural Network Training with Communication Reduction
by: Zhang, Shuai, et al.
Published: (2024)
by: Zhang, Shuai, et al.
Published: (2024)
Robust Federated Finetuning of Foundation Models via Alternating Minimization of LoRA
by: Chen, Shuangyi, et al.
Published: (2024)
by: Chen, Shuangyi, et al.
Published: (2024)
FedRFQ: Prototype-Based Federated Learning with Reduced Redundancy, Minimal Failure, and Enhanced Quality
by: Yan, Biwei, et al.
Published: (2024)
by: Yan, Biwei, et al.
Published: (2024)
Locally Estimated Global Perturbations are Better than Local Perturbations for Federated Sharpness-aware Minimization
by: Fan, Ziqing, et al.
Published: (2024)
by: Fan, Ziqing, et al.
Published: (2024)
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
by: Sui, Yifan, et al.
Published: (2025)
by: Sui, Yifan, et al.
Published: (2025)
Similar Items
-
Communication Efficient Distributed Training with Distributed Lion
by: Liu, Bo, et al.
Published: (2024) -
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
by: Li, Shigang, et al.
Published: (2020) -
Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
by: Li, Shigang, et al.
Published: (2019) -
Lion: Minimizing Distributed Transactions through Adaptive Replica Provision (Extended Version)
by: Zheng, Qiushi, et al.
Published: (2024) -
Accelerating Large Language Model Training with 4D Parallelism and Memory Consumption Estimator
by: Fujii, Kazuki, et al.
Published: (2024)