Dynamic Hierarchical Birkhoff-von Neumann Decomposition for All-to-All GPU Communication
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Yen-Chieh, Chang, Cheng-Shang, Lee, Duan-Shin, Chao, H. Jonathan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FAST: An Efficient Scheduler for All-to-All GPU Communication
por: Lei, Yiran, et al.
Publicado: (2025)
por: Lei, Yiran, et al.
Publicado: (2025)
AllReduce Scheduling with Hierarchical Deep Reinforcement Learning
por: Wei, Yufan, et al.
Publicado: (2025)
por: Wei, Yufan, et al.
Publicado: (2025)
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
por: Xu, Heng, et al.
Publicado: (2025)
por: Xu, Heng, et al.
Publicado: (2025)
Efficient All-to-All Collective Communication Schedules for Direct-Connect Topologies
por: Basu, Prithwish, et al.
Publicado: (2023)
por: Basu, Prithwish, et al.
Publicado: (2023)
Revisiting Bruck: Phase-Efficient All-to-All Communication in Reconfigurable Networks
por: Juerss, Anton, et al.
Publicado: (2026)
por: Juerss, Anton, et al.
Publicado: (2026)
Short-circuiting Rings for Low-Latency AllReduce
por: Hammer, Sarah-Michelle, et al.
Publicado: (2025)
por: Hammer, Sarah-Michelle, et al.
Publicado: (2025)
Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
por: Juerss, Anton, et al.
Publicado: (2026)
por: Juerss, Anton, et al.
Publicado: (2026)
Fast Multichannel Topology Discovery in Cognitive Radio Networks
por: Wang, Yung-Li, et al.
Publicado: (2025)
por: Wang, Yung-Li, et al.
Publicado: (2025)
OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the Cloud
por: Warraich, Ertza, et al.
Publicado: (2023)
por: Warraich, Ertza, et al.
Publicado: (2023)
An Online Fragmentation-Aware GPU Scheduler for Multi-Tenant MIG-based Clouds
por: Zambianco, Marco, et al.
Publicado: (2025)
por: Zambianco, Marco, et al.
Publicado: (2025)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
por: Du, Chengze, et al.
Publicado: (2025)
por: Du, Chengze, et al.
Publicado: (2025)
From Skew to Symmetry: Node-Interconnect Multi-Path Balancing with Execution-time Planning for Modern GPU Clusters
por: Yao, Jinghan, et al.
Publicado: (2026)
por: Yao, Jinghan, et al.
Publicado: (2026)
A Survey on Resource Management in Joint Communication and Computing-Embedded SAGIN
por: Chen, Qian, et al.
Publicado: (2024)
por: Chen, Qian, et al.
Publicado: (2024)
A Task Decomposition and Planning Framework for Efficient LLM Inference in AI-Enabled WiFi-Offload Networks
por: Han, Mingqi, et al.
Publicado: (2026)
por: Han, Mingqi, et al.
Publicado: (2026)
Towards Integrated Energy-Communication-Transportation Hub: A Base-Station-Centric Design in 5G and Beyond
por: Shen, Linfeng, et al.
Publicado: (2025)
por: Shen, Linfeng, et al.
Publicado: (2025)
Diving into 3D Parallelism with Heterogeneous Spot Instance GPUs: Design and Implications
por: Wang, Yuxiao, et al.
Publicado: (2025)
por: Wang, Yuxiao, et al.
Publicado: (2025)
Accelerating Stable Matching between Workers and Spatial-Temporal Tasks for Dynamic MCS: A Stagewise Service Trading Approach
por: Qi, Houyi, et al.
Publicado: (2025)
por: Qi, Houyi, et al.
Publicado: (2025)
Topology-aware Microservice Architecture in Edge Networks: Deployment Optimization and Implementation
por: Chen, Yuang, et al.
Publicado: (2025)
por: Chen, Yuang, et al.
Publicado: (2025)
Don't Let a Few Network Failures Slow the Entire AllReduce
por: Chen, Peiqing, et al.
Publicado: (2026)
por: Chen, Peiqing, et al.
Publicado: (2026)
Rina: Enhancing Ring-AllReduce with In-network Aggregation in Distributed Model Training
por: Chen, Zixuan, et al.
Publicado: (2024)
por: Chen, Zixuan, et al.
Publicado: (2024)
Enabling Scalability in Asynchronous and Bidirectional Communication in LPWAN
por: Rahman, Mahbubur
Publicado: (2025)
por: Rahman, Mahbubur
Publicado: (2025)
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services
por: Tang, Lingfeng, et al.
Publicado: (2025)
por: Tang, Lingfeng, et al.
Publicado: (2025)
Harvest: Adaptive Photonic Switching Schedules for Collective Communication in Scale-up Domains
por: Rahman, Mahir, et al.
Publicado: (2026)
por: Rahman, Mahir, et al.
Publicado: (2026)
LOAM: Low-latency Communication, Caching, and Computation Placement in Data-Intensive Computing Networks
por: Zhang, Jinkun, et al.
Publicado: (2024)
por: Zhang, Jinkun, et al.
Publicado: (2024)
Self-Healing Network of Interconnected Edge Devices Empowered by Infrastructure-as-Code and LoRa Communication
por: Carson, Rob, et al.
Publicado: (2025)
por: Carson, Rob, et al.
Publicado: (2025)
Meili: Enabling SmartNIC as a Service in the Cloud
por: Su, Qiang, et al.
Publicado: (2023)
por: Su, Qiang, et al.
Publicado: (2023)
OrchestrRL: Dynamic Compute and Network Orchestration for Disaggregated RL
por: Tan, Xin, et al.
Publicado: (2026)
por: Tan, Xin, et al.
Publicado: (2026)
EdgeLoc: A Communication-Adaptive Parallel System for Real-Time Localization in Infrastructure-Assisted Autonomous Driving
por: Liu, Boyi, et al.
Publicado: (2024)
por: Liu, Boyi, et al.
Publicado: (2024)
Bodyless Block Propagation: TPS Fully Scalable Blockchain with Pre-Validation
por: Zhao, Chonghe, et al.
Publicado: (2022)
por: Zhao, Chonghe, et al.
Publicado: (2022)
EdgeLinker: Practical Blockchain-based Framework for Healthcare Fog Applications to Enhance Security in Edge-IoT Data Communications
por: Zarkesh, Mahdi Akbari, et al.
Publicado: (2024)
por: Zarkesh, Mahdi Akbari, et al.
Publicado: (2024)
A Uniqueness Theorem for Distributed Computation under Physical Constraint
por: Ren, Zhiyuan, et al.
Publicado: (2025)
por: Ren, Zhiyuan, et al.
Publicado: (2025)
Hurry: Dynamic Collaborative Framework For Low-orbit Mega-Constellation Data Downloading
por: Luo, Handong, et al.
Publicado: (2024)
por: Luo, Handong, et al.
Publicado: (2024)
Dynamic DAG-Application Scheduling for Multi-Tier Edge Computing in Heterogeneous Networks
por: Li, Xiang, et al.
Publicado: (2024)
por: Li, Xiang, et al.
Publicado: (2024)
Dynamic Edge Server Selection in Time-Varying Environments: A Reliability-Aware Predictive Approach
por: Burbano, Jaime Sebastian, et al.
Publicado: (2025)
por: Burbano, Jaime Sebastian, et al.
Publicado: (2025)
PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services
por: Yang, Zheming, et al.
Publicado: (2024)
por: Yang, Zheming, et al.
Publicado: (2024)
Reactive Orchestration for Hierarchical Federated Learning Under a Communication Cost Budget
por: Čilić, Ivan, et al.
Publicado: (2024)
por: Čilić, Ivan, et al.
Publicado: (2024)
DRDST: Low-latency DAG Consensus through Robust Dynamic Sharding and Tree-broadcasting for IoV
por: Chen, Runhua, et al.
Publicado: (2024)
por: Chen, Runhua, et al.
Publicado: (2024)
Performance Analysis of Internet of Vehicles Mesh Networks Based on Actual Switch Models
por: Hu, Jialin, et al.
Publicado: (2024)
por: Hu, Jialin, et al.
Publicado: (2024)
SAKURAONE: An Open Ethernet-Based AI HPC System and Its Observed Workload Dynamics in a Single-Tenant LLM Development Environment
por: Konishi, Fumikazu, et al.
Publicado: (2026)
por: Konishi, Fumikazu, et al.
Publicado: (2026)
Earth+: on-board satellite imagery compression leveraging historical earth observations
por: Du, Kuntai, et al.
Publicado: (2024)
por: Du, Kuntai, et al.
Publicado: (2024)
Ejemplares similares
-
FAST: An Efficient Scheduler for All-to-All GPU Communication
por: Lei, Yiran, et al.
Publicado: (2025) -
AllReduce Scheduling with Hierarchical Deep Reinforcement Learning
por: Wei, Yufan, et al.
Publicado: (2025) -
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
por: Xu, Heng, et al.
Publicado: (2025) -
Efficient All-to-All Collective Communication Schedules for Direct-Connect Topologies
por: Basu, Prithwish, et al.
Publicado: (2023) -
Revisiting Bruck: Phase-Efficient All-to-All Communication in Reconfigurable Networks
por: Juerss, Anton, et al.
Publicado: (2026)