Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
Fuente:
arXiv
Saved in:
| Main Authors: | Papavasileiou, Ioannis, Prabhakar, Sairam, Deo, Indu Kant, Makovejs, Sergejs |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Multi-Port Concurrent Communication Model for handling Compute Intensive Tasks on Distributed Satellite System Constellations
by: Veeravalli, Bharadwaj
Published: (2026)
by: Veeravalli, Bharadwaj
Published: (2026)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
by: Vatsavai, Sairam Sri, et al.
Published: (2025)
by: Vatsavai, Sairam Sri, et al.
Published: (2025)
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
by: Jeffery, Andrew, et al.
Published: (2024)
by: Jeffery, Andrew, et al.
Published: (2024)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
by: McDonald, Jesse, et al.
Published: (2024)
by: McDonald, Jesse, et al.
Published: (2024)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
by: Zhuang, Chen, et al.
Published: (2025)
by: Zhuang, Chen, et al.
Published: (2025)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
by: Karfakis, George, et al.
Published: (2025)
by: Karfakis, George, et al.
Published: (2025)
PlantD: Performance, Latency ANalysis, and Testing for Data Pipelines -- An Open Source Measurement, Testing, and Simulation Framework
by: Bogart, Christopher, et al.
Published: (2025)
by: Bogart, Christopher, et al.
Published: (2025)
ISO: Overlap of Computation and Communication within Seqenence For LLM Inference
by: Xiao, Bin, et al.
Published: (2024)
by: Xiao, Bin, et al.
Published: (2024)
Energy-Aware Computing in the Year 2026
by: Tchakoute, Roblex Nana, et al.
Published: (2026)
by: Tchakoute, Roblex Nana, et al.
Published: (2026)
On Orchestrating Parallel Broadcasts for Distributed Ledgers
by: Sheng, Peiyao, et al.
Published: (2024)
by: Sheng, Peiyao, et al.
Published: (2024)
An Online Probabilistic Distributed Tracing System
by: Toslali, M., et al.
Published: (2024)
by: Toslali, M., et al.
Published: (2024)
Towards a Peer-to-Peer Data Distribution Layer for Efficient and Collaborative Resource Optimization of Distributed Dataflow Applications
by: Scheinert, Dominik, et al.
Published: (2023)
by: Scheinert, Dominik, et al.
Published: (2023)
Hiku: Pull-Based Scheduling for Serverless Computing
by: Akbari, Saman, et al.
Published: (2025)
by: Akbari, Saman, et al.
Published: (2025)
Optimal Configuration of API Resources in Cloud Native Computing
by: Truyen, Eddy, et al.
Published: (2025)
by: Truyen, Eddy, et al.
Published: (2025)
Kubernetes in Action: Exploring the Performance of Kubernetes Distributions in the Cloud
by: Aqasizade, Hossein, et al.
Published: (2024)
by: Aqasizade, Hossein, et al.
Published: (2024)
Performance Impact of Containerized METADOCK 2 on Heterogeneous Platforms
by: Banegas-Luna, Antonio Jesús, et al.
Published: (2025)
by: Banegas-Luna, Antonio Jesús, et al.
Published: (2025)
Ridgeline: A 2D Roofline Model for Distributed Systems
by: Checconi, Fabio, et al.
Published: (2022)
by: Checconi, Fabio, et al.
Published: (2022)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
by: Lacey, Dane C., et al.
Published: (2024)
by: Lacey, Dane C., et al.
Published: (2024)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
by: Zhuang, Chen, et al.
Published: (2024)
by: Zhuang, Chen, et al.
Published: (2024)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
by: Wang, Yuxin, et al.
Published: (2023)
by: Wang, Yuxin, et al.
Published: (2023)
Communication-Aware Diffusion Load Balancing for Persistently Interacting Objects
by: Taylor, Maya, et al.
Published: (2026)
by: Taylor, Maya, et al.
Published: (2026)
SProBench: Stream Processing Benchmark for High Performance Computing Infrastructure
by: Kulkarni, Apurv Deepak, et al.
Published: (2025)
by: Kulkarni, Apurv Deepak, et al.
Published: (2025)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
by: Ramesh, Risshab Srinivas
Published: (2024)
by: Ramesh, Risshab Srinivas
Published: (2024)
Optimizations on Graph-Level for Domain Specific Computations in Julia and Application to QED
by: Reinhard, Anton, et al.
Published: (2025)
by: Reinhard, Anton, et al.
Published: (2025)
Universal Workers: A Vision for Eliminating Cold Starts in Serverless Computing
by: Akbari, Saman, et al.
Published: (2025)
by: Akbari, Saman, et al.
Published: (2025)
Modeling the Effect of Data Redundancy on Speedup in MLFMA Near-Field Computation
by: Sadeghi, Morteza
Published: (2025)
by: Sadeghi, Morteza
Published: (2025)
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
by: Shan, Baodi, et al.
Published: (2024)
by: Shan, Baodi, et al.
Published: (2024)
Resource Management Schemes for Cloud-Native Platforms with Computing Containers of Docker and Kubernetes
by: Mao, Ying, et al.
Published: (2020)
by: Mao, Ying, et al.
Published: (2020)
Cost-Performance Evaluation of General Compute Instances: AWS, Azure, GCP, and OCI
by: Tharwani, Jay, et al.
Published: (2024)
by: Tharwani, Jay, et al.
Published: (2024)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
by: Xu, Jingwei, et al.
Published: (2025)
by: Xu, Jingwei, et al.
Published: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
by: Lin, Mao, et al.
Published: (2026)
by: Lin, Mao, et al.
Published: (2026)
CASA: A Framework for SLO and Carbon-Aware Autoscaling and Scheduling in Serverless Cloud Computing
by: Qi, S., et al.
Published: (2024)
by: Qi, S., et al.
Published: (2024)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
by: Rashid, Md Hasanur, et al.
Published: (2026)
by: Rashid, Md Hasanur, et al.
Published: (2026)
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
by: Andersson, Måns I., et al.
Published: (2025)
by: Andersson, Måns I., et al.
Published: (2025)
DREAMS: Decentralized Resource Allocation and Service Management across the Compute Continuum Using Service Affinity
by: Dinh-Tuan, Hai, et al.
Published: (2025)
by: Dinh-Tuan, Hai, et al.
Published: (2025)
The SAP Cloud Infrastructure Dataset: A Reality Check of Scheduling and Placement of VMs in Cloud Computing
by: Uhlig, Arno, et al.
Published: (2025)
by: Uhlig, Arno, et al.
Published: (2025)
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
by: Lurati, Milo, et al.
Published: (2024)
by: Lurati, Milo, et al.
Published: (2024)
On the Partitioning of GPU Power among Multi-Instances
by: Vamja, Tirth, et al.
Published: (2025)
by: Vamja, Tirth, et al.
Published: (2025)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
by: Zhang, Yaozheng, et al.
Published: (2025)
by: Zhang, Yaozheng, et al.
Published: (2025)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
by: Ng, Nathan, et al.
Published: (2026)
by: Ng, Nathan, et al.
Published: (2026)
Similar Items
-
A Multi-Port Concurrent Communication Model for handling Compute Intensive Tasks on Distributed Satellite System Constellations
by: Veeravalli, Bharadwaj
Published: (2026) -
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
by: Vatsavai, Sairam Sri, et al.
Published: (2025) -
Reducing Tail Latencies Through Environment- and Neighbour-aware Thread Management
by: Jeffery, Andrew, et al.
Published: (2024) -
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
by: McDonald, Jesse, et al.
Published: (2024) -
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
by: Zhuang, Chen, et al.
Published: (2025)