From Skew to Symmetry: Node-Interconnect Multi-Path Balancing with Execution-time Planning for Modern GPU Clusters
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Jinghan, Kandadi, Kaushik, Ramesh, Bharath, Subramoni, Hari, Panda, Dhabaleswar K. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Extreme-Scale Interconnection Networks
by: Cano, Alejandro, et al.
Published: (2026)
by: Cano, Alejandro, et al.
Published: (2026)
Characterizing Communication Patterns in Distributed Large Language Model Inference
by: Xu, Lang, et al.
Published: (2025)
by: Xu, Lang, et al.
Published: (2025)
NET4EXA: Pioneering the Future of Interconnects for Supercomputing and AI
by: Martinelli, Michele, et al.
Published: (2026)
by: Martinelli, Michele, et al.
Published: (2026)
LACIN: Linearly Arranged Complete Interconnection Networks
by: Beivide, Ramón, et al.
Published: (2026)
by: Beivide, Ramón, et al.
Published: (2026)
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
by: Yao, Jinghan, et al.
Published: (2024)
by: Yao, Jinghan, et al.
Published: (2024)
FlowTracer: A Tool for Uncovering Network Path Usage Imbalance in AI Training Clusters
by: Jamil, Hasibul, et al.
Published: (2024)
by: Jamil, Hasibul, et al.
Published: (2024)
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services
by: Tang, Lingfeng, et al.
Published: (2025)
by: Tang, Lingfeng, et al.
Published: (2025)
Self-Healing Network of Interconnected Edge Devices Empowered by Infrastructure-as-Code and LoRa Communication
by: Carson, Rob, et al.
Published: (2025)
by: Carson, Rob, et al.
Published: (2025)
FAST: An Efficient Scheduler for All-to-All GPU Communication
by: Lei, Yiran, et al.
Published: (2025)
by: Lei, Yiran, et al.
Published: (2025)
Efficient Fog Node Placement using Nature-Inspired Metaheuristic for IoT Applications
by: Naouri, Abdenacer, et al.
Published: (2023)
by: Naouri, Abdenacer, et al.
Published: (2023)
Accelerating Time-to-Science by Streaming Detector Data Directly into Perlmutter Compute Nodes
by: Welborn, Samuel S., et al.
Published: (2024)
by: Welborn, Samuel S., et al.
Published: (2024)
Dynamic Hierarchical Birkhoff-von Neumann Decomposition for All-to-All GPU Communication
by: Wu, Yen-Chieh, et al.
Published: (2026)
by: Wu, Yen-Chieh, et al.
Published: (2026)
An Online Fragmentation-Aware GPU Scheduler for Multi-Tenant MIG-based Clouds
by: Zambianco, Marco, et al.
Published: (2025)
by: Zambianco, Marco, et al.
Published: (2025)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
by: Du, Chengze, et al.
Published: (2025)
by: Du, Chengze, et al.
Published: (2025)
Selecting Relay Nodes Based on Evaluation Results to Enhance P2P Broadcasting Efficiency
by: Huang, Chunlin
Published: (2024)
by: Huang, Chunlin
Published: (2024)
EvalNet: A Practical Toolchain for Generation and Analysis of Extreme-Scale Interconnects
by: Besta, Maciej, et al.
Published: (2021)
by: Besta, Maciej, et al.
Published: (2021)
Deduplicator: When Computation Reuse Meets Load Balancing at the Network Edge
by: Azad, Md Washik Al, et al.
Published: (2024)
by: Azad, Md Washik Al, et al.
Published: (2024)
CRAFT: Latency and Cost-Aware Genetic-Based Framework for Node Placement in Edge-Fog Environments
by: Mahdizadeh, Soheil, et al.
Published: (2025)
by: Mahdizadeh, Soheil, et al.
Published: (2025)
Whack-a-Mole: Deterministic Packet Spraying Across Multiple Network Paths
by: Luby, Michael, et al.
Published: (2025)
by: Luby, Michael, et al.
Published: (2025)
Optimal Oblivious Load-Balancing for Sparse Traffic in Large-Scale Satellite Networks
by: Ramakanth, Rudrapatna Vallabh, et al.
Published: (2026)
by: Ramakanth, Rudrapatna Vallabh, et al.
Published: (2026)
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
by: Xu, Heng, et al.
Published: (2025)
by: Xu, Heng, et al.
Published: (2025)
QoS-Aware Load Balancing in the Computing Continuum via Multi-Player Bandits
by: Čilić, Ivan, et al.
Published: (2025)
by: Čilić, Ivan, et al.
Published: (2025)
Ethereal: Divide and Conquer Network Load Balancing in Large-Scale Distributed Training
by: Addanki, Vamsi, et al.
Published: (2024)
by: Addanki, Vamsi, et al.
Published: (2024)
Design and Operation of Shared Machine Learning Clusters on Campus
by: Xu, Kaiqiang, et al.
Published: (2021)
by: Xu, Kaiqiang, et al.
Published: (2021)
Heterogeneity-aware P2P Wireless Energy Transfer for Balanced Energy Distribution
by: Ojha, Tamoghna, et al.
Published: (2022)
by: Ojha, Tamoghna, et al.
Published: (2022)
cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
Intent-driven Diffusion-based Path for Mobile Data Collector in IoT-enabled Dense WSNs
by: Boda, Uma Mahesh, et al.
Published: (2026)
by: Boda, Uma Mahesh, et al.
Published: (2026)
Toward Co-adapting Machine Learning Job Shape and Cluster Topology
by: Chen, Shawn Shuoshuo, et al.
Published: (2025)
by: Chen, Shawn Shuoshuo, et al.
Published: (2025)
YUHENG-OS: A Cloud-Native Space Cluster Operating System
by: Zhang, Jin, et al.
Published: (2026)
by: Zhang, Jin, et al.
Published: (2026)
GORGO: Maximizing KV-Cache Reuse While Minimizing Network Latency in Cross-Region LLM Load Balancing
by: Toniolo, Alessio Ricci, et al.
Published: (2026)
by: Toniolo, Alessio Ricci, et al.
Published: (2026)
ClusterSlice: A Zero-touch Deployment Platform for the Edge Cloud Continuum
by: Mamatas, Lefteris, et al.
Published: (2024)
by: Mamatas, Lefteris, et al.
Published: (2024)
LIDC: A Location Independent Multi-Cluster Computing Framework for Data Intensive Science
by: Timilsina, Sankalpa, et al.
Published: (2025)
by: Timilsina, Sankalpa, et al.
Published: (2025)
A New Classification of Clustering-based for Different Problems in Different Wireless Ad-hoc Networks
by: Boualem, Adda, et al.
Published: (2024)
by: Boualem, Adda, et al.
Published: (2024)
A Combined Environmental Monitoring Framework based on WSN Clustering and VANET Edge Computation Offloading
by: Mamalis, Basilis, et al.
Published: (2024)
by: Mamalis, Basilis, et al.
Published: (2024)
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
by: Yao, Jinghan, et al.
Published: (2024)
by: Yao, Jinghan, et al.
Published: (2024)
A Task Decomposition and Planning Framework for Efficient LLM Inference in AI-Enabled WiFi-Offload Networks
by: Han, Mingqi, et al.
Published: (2026)
by: Han, Mingqi, et al.
Published: (2026)
Accelerating Large Language Model Training with Hybrid GPU-based Compression
by: Xu, Lang, et al.
Published: (2024)
by: Xu, Lang, et al.
Published: (2024)
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning
by: Xu, Lang, et al.
Published: (2025)
by: Xu, Lang, et al.
Published: (2025)
Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics
by: Ma, Bole, et al.
Published: (2026)
by: Ma, Bole, et al.
Published: (2026)
Collective Communication Profiling of Modern-day Machine Learning Workloads
by: Gupta, Jit, et al.
Published: (2025)
by: Gupta, Jit, et al.
Published: (2025)
Similar Items
-
Extreme-Scale Interconnection Networks
by: Cano, Alejandro, et al.
Published: (2026) -
Characterizing Communication Patterns in Distributed Large Language Model Inference
by: Xu, Lang, et al.
Published: (2025) -
NET4EXA: Pioneering the Future of Interconnects for Supercomputing and AI
by: Martinelli, Michele, et al.
Published: (2026) -
LACIN: Linearly Arranged Complete Interconnection Networks
by: Beivide, Ramón, et al.
Published: (2026) -
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
by: Yao, Jinghan, et al.
Published: (2024)