Improving training time and GPU utilization in geo-distributed language model training
Fuente:
arXiv
Saved in:
| Main Authors: | Palak, Reddy, Tella Rajashekhar, Kataria, Bhaskar, Gandhi, Rohan, Tandon, Karan, Bhattacherjee, Debopam, Padmanabhan, Venkata N. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BeLLMan: Controlling LLM Congestion
by: Reddy, Tella Rajashekhar, et al.
Published: (2025)
by: Reddy, Tella Rajashekhar, et al.
Published: (2025)
AI Greenferencing: Routing AI Inferencing to Green Modular Data Centers with Heron
by: Reddy, Tella Rajashekhar, et al.
Published: (2025)
by: Reddy, Tella Rajashekhar, et al.
Published: (2025)
XWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy Farms
by: Reddy, Tella Rajashekhar, et al.
Published: (2026)
by: Reddy, Tella Rajashekhar, et al.
Published: (2026)
KnapsackLB: Enabling Performance-Aware Layer-4 Load Balancing
by: Gandhi, Rohan, et al.
Published: (2024)
by: Gandhi, Rohan, et al.
Published: (2024)
Fog enabled distributed training architecture for federated learning
by: Kumar, Aditya, et al.
Published: (2024)
by: Kumar, Aditya, et al.
Published: (2024)
Tetris: Efficient Intra-Datacenter Calls Packing for Large Conferencing Services
by: Gandhi, Rohan, et al.
Published: (2025)
by: Gandhi, Rohan, et al.
Published: (2025)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
by: Wang, Tianyu, et al.
Published: (2024)
by: Wang, Tianyu, et al.
Published: (2024)
Simulating LLM training workloads for heterogeneous compute and network infrastructure
by: Kumar, Sumit, et al.
Published: (2025)
by: Kumar, Sumit, et al.
Published: (2025)
Beyond Microservices: Testing Web-Scale RCA Methods on GPU-Driven LLM Workloads
by: Scheinert, Dominik, et al.
Published: (2026)
by: Scheinert, Dominik, et al.
Published: (2026)
Sparse Checkpointing for Fast and Reliable MoE Training
by: Gandhi, Swapnil, et al.
Published: (2024)
by: Gandhi, Swapnil, et al.
Published: (2024)
PISA: An Adversarial Approach To Comparing Task Graph Scheduling Algorithms
by: Coleman, Jared, et al.
Published: (2024)
by: Coleman, Jared, et al.
Published: (2024)
Performance-Driven Optimization of Parallel Breadth-First Search
by: Bhaskar, Marati, et al.
Published: (2025)
by: Bhaskar, Marati, et al.
Published: (2025)
HeLoCo: Efficient asynchronous low-communication training under data and device heterogeneity
by: Asif, Abdullah Al, et al.
Published: (2026)
by: Asif, Abdullah Al, et al.
Published: (2026)
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
Accelerating Biclique Counting on GPU
by: Qiu, Linshan, et al.
Published: (2024)
by: Qiu, Linshan, et al.
Published: (2024)
GPU Sharing with Triples Mode
by: Byun, Chansup, et al.
Published: (2024)
by: Byun, Chansup, et al.
Published: (2024)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
by: Lee, Munkyu, et al.
Published: (2024)
by: Lee, Munkyu, et al.
Published: (2024)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
by: Sojoodi, Amirhossein, et al.
Published: (2026)
by: Sojoodi, Amirhossein, et al.
Published: (2026)
DuaLip-GPU Technical Report
by: Dexter, Gregory, et al.
Published: (2026)
by: Dexter, Gregory, et al.
Published: (2026)
Incidence Constraints in Hypergraph Partitioning on GPU
by: Ronzani, Marco, et al.
Published: (2026)
by: Ronzani, Marco, et al.
Published: (2026)
Predictable LLM Serving on GPU Clusters
by: Darzi, Erfan, et al.
Published: (2025)
by: Darzi, Erfan, et al.
Published: (2025)
GPU Accelerated Sparse Cholesky Factorization
by: Karsavuran, M. Ozan, et al.
Published: (2024)
by: Karsavuran, M. Ozan, et al.
Published: (2024)
Heat: Satellite's meat is GPU's poison
by: Yuan, Zhehu, et al.
Published: (2024)
by: Yuan, Zhehu, et al.
Published: (2024)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
by: Gu, Jianfeng, et al.
Published: (2025)
by: Gu, Jianfeng, et al.
Published: (2025)
Optimizing Bloom Filters for Modern GPU Architectures
by: Jünger, Daniel, et al.
Published: (2025)
by: Jünger, Daniel, et al.
Published: (2025)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
by: Stoyanov, Radostin, et al.
Published: (2025)
by: Stoyanov, Radostin, et al.
Published: (2025)
Understanding the Landscape of Ampere GPU Memory Errors
by: Zhu, Zhu, et al.
Published: (2025)
by: Zhu, Zhu, et al.
Published: (2025)
Methodology for GPU Frequency Switching Latency Measurement
by: Velicka, Daniel, et al.
Published: (2025)
by: Velicka, Daniel, et al.
Published: (2025)
Adaptive Multidimensional Quadrature on Multi-GPU Systems
by: Tonarelli, Melanie, et al.
Published: (2025)
by: Tonarelli, Melanie, et al.
Published: (2025)
GPU-Accelerated Batch-Dynamic Subgraph Matching
by: Qiu, Linshan, et al.
Published: (2024)
by: Qiu, Linshan, et al.
Published: (2024)
A Multi-Objective Framework for Optimizing GPU-Enabled VM Placement in Cloud Data Centers with Multi-Instance GPU Technology
by: Siavashi, Ahmad, et al.
Published: (2025)
by: Siavashi, Ahmad, et al.
Published: (2025)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
by: Maurya, Avinash, et al.
Published: (2024)
by: Maurya, Avinash, et al.
Published: (2024)
Zen-Attention: A Compiler Framework for Dynamic Attention Folding on AMD NPUs
by: Deshmukh, Aadesh, et al.
Published: (2025)
by: Deshmukh, Aadesh, et al.
Published: (2025)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
Understanding GPU Triggering APIs for MPI+X Communication
by: Bridges, Patrick G., et al.
Published: (2024)
by: Bridges, Patrick G., et al.
Published: (2024)
Optimal Resource Efficiency with Fairness in Heterogeneous GPU Clusters
by: Mo, Zizhao, et al.
Published: (2024)
by: Mo, Zizhao, et al.
Published: (2024)
DeepOps & SLURM: Your GPU Cluster Guide
by: Majee, Arindam
Published: (2024)
by: Majee, Arindam
Published: (2024)
Characterization-Guided GPU Fault Resilience in NVIDIA MPS
by: Liu, Rixin, et al.
Published: (2026)
by: Liu, Rixin, et al.
Published: (2026)
PICO: Accelerating All k-Core Paradigms on GPU
by: Zhao, Chen, et al.
Published: (2024)
by: Zhao, Chen, et al.
Published: (2024)
AGILE: Lightweight and Efficient Asynchronous GPU-SSD Integration
by: Yang, Zhuoping, et al.
Published: (2025)
by: Yang, Zhuoping, et al.
Published: (2025)
Similar Items
-
BeLLMan: Controlling LLM Congestion
by: Reddy, Tella Rajashekhar, et al.
Published: (2025) -
AI Greenferencing: Routing AI Inferencing to Green Modular Data Centers with Heron
by: Reddy, Tella Rajashekhar, et al.
Published: (2025) -
XWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy Farms
by: Reddy, Tella Rajashekhar, et al.
Published: (2026) -
KnapsackLB: Enabling Performance-Aware Layer-4 Load Balancing
by: Gandhi, Rohan, et al.
Published: (2024) -
Fog enabled distributed training architecture for federated learning
by: Kumar, Aditya, et al.
Published: (2024)