AllReduce Scheduling with Hierarchical Deep Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Yufan, Liu, Mickel, Wu, Wenfei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the Cloud
by: Warraich, Ertza, et al.
Published: (2023)
by: Warraich, Ertza, et al.
Published: (2023)
Short-circuiting Rings for Low-Latency AllReduce
by: Hammer, Sarah-Michelle, et al.
Published: (2025)
by: Hammer, Sarah-Michelle, et al.
Published: (2025)
Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
by: Juerss, Anton, et al.
Published: (2026)
by: Juerss, Anton, et al.
Published: (2026)
Don't Let a Few Network Failures Slow the Entire AllReduce
by: Chen, Peiqing, et al.
Published: (2026)
by: Chen, Peiqing, et al.
Published: (2026)
Rina: Enhancing Ring-AllReduce with In-network Aggregation in Distributed Model Training
by: Chen, Zixuan, et al.
Published: (2024)
by: Chen, Zixuan, et al.
Published: (2024)
EdgeTimer: Adaptive Multi-Timescale Scheduling in Mobile Edge Computing with Deep Reinforcement Learning
by: Hao, Yijun, et al.
Published: (2024)
by: Hao, Yijun, et al.
Published: (2024)
FAST: An Efficient Scheduler for All-to-All GPU Communication
by: Lei, Yiran, et al.
Published: (2025)
by: Lei, Yiran, et al.
Published: (2025)
Efficient All-to-All Collective Communication Schedules for Direct-Connect Topologies
by: Basu, Prithwish, et al.
Published: (2023)
by: Basu, Prithwish, et al.
Published: (2023)
Dynamic Hierarchical Birkhoff-von Neumann Decomposition for All-to-All GPU Communication
by: Wu, Yen-Chieh, et al.
Published: (2026)
by: Wu, Yen-Chieh, et al.
Published: (2026)
EdgeLoc: A Communication-Adaptive Parallel System for Real-Time Localization in Infrastructure-Assisted Autonomous Driving
by: Liu, Boyi, et al.
Published: (2024)
by: Liu, Boyi, et al.
Published: (2024)
Multi-stage Flow Scheduling for LLM Serving
by: Sun, Yijun, et al.
Published: (2026)
by: Sun, Yijun, et al.
Published: (2026)
Revisiting Bruck: Phase-Efficient All-to-All Communication in Reconfigurable Networks
by: Juerss, Anton, et al.
Published: (2026)
by: Juerss, Anton, et al.
Published: (2026)
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
by: Xu, Heng, et al.
Published: (2025)
by: Xu, Heng, et al.
Published: (2025)
Hierarchical Online-Scheduling for Energy-Efficient Split Inference with Progressive Transmission
by: Tang, Zengzipeng, et al.
Published: (2026)
by: Tang, Zengzipeng, et al.
Published: (2026)
Carbon-Aware Temporal Data Transfer Scheduling Across Cloud Datacenters
by: Rodrigues, Elvis, et al.
Published: (2025)
by: Rodrigues, Elvis, et al.
Published: (2025)
Multi-Source Coflow Scheduling in Collaborative Edge Computing with Multihop Network
by: Sahni, Yuvraj, et al.
Published: (2024)
by: Sahni, Yuvraj, et al.
Published: (2024)
An Online Fragmentation-Aware GPU Scheduler for Multi-Tenant MIG-based Clouds
by: Zambianco, Marco, et al.
Published: (2025)
by: Zambianco, Marco, et al.
Published: (2025)
Dynamic DAG-Application Scheduling for Multi-Tier Edge Computing in Heterogeneous Networks
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Harvest: Adaptive Photonic Switching Schedules for Collective Communication in Scale-up Domains
by: Rahman, Mahir, et al.
Published: (2026)
by: Rahman, Mahir, et al.
Published: (2026)
PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services
by: Yang, Zheming, et al.
Published: (2024)
by: Yang, Zheming, et al.
Published: (2024)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
by: Du, Chengze, et al.
Published: (2025)
by: Du, Chengze, et al.
Published: (2025)
Q-adaptive: A Multi-Agent Reinforcement Learning Based Routing on Dragonfly Network
by: Kang, Yao, et al.
Published: (2024)
by: Kang, Yao, et al.
Published: (2024)
Real-Time Scheduling for 802.1Qbv Time-Sensitive Networking (TSN): A Systematic Review and Experimental Study
by: Xue, Chuanyu, et al.
Published: (2023)
by: Xue, Chuanyu, et al.
Published: (2023)
Towards Practical Overlay Networks for Decentralized Federated Learning
by: Hua, Yifan, et al.
Published: (2024)
by: Hua, Yifan, et al.
Published: (2024)
DeepEdge: A Deep Reinforcement Learning based Task Orchestrator for Edge Computing
by: Yamansavascilar, Baris, et al.
Published: (2021)
by: Yamansavascilar, Baris, et al.
Published: (2021)
Towards Integrated Energy-Communication-Transportation Hub: A Base-Station-Centric Design in 5G and Beyond
by: Shen, Linfeng, et al.
Published: (2025)
by: Shen, Linfeng, et al.
Published: (2025)
Resource Allocation Driven by Large Models in Future Semantic-Aware Networks
by: Zhang, Haijun, et al.
Published: (2025)
by: Zhang, Haijun, et al.
Published: (2025)
TraDE: Network and Traffic-aware Adaptive Scheduling for Microservices Under Dynamics
by: Chen, Ming, et al.
Published: (2024)
by: Chen, Ming, et al.
Published: (2024)
Diving into 3D Parallelism with Heterogeneous Spot Instance GPUs: Design and Implications
by: Wang, Yuxiao, et al.
Published: (2025)
by: Wang, Yuxiao, et al.
Published: (2025)
Recursive Offloading for LLM Serving in Multi-tier Networks
by: Wu, Zhiyuan, et al.
Published: (2025)
by: Wu, Zhiyuan, et al.
Published: (2025)
Meili: Enabling SmartNIC as a Service in the Cloud
by: Su, Qiang, et al.
Published: (2023)
by: Su, Qiang, et al.
Published: (2023)
Design and Operation of Shared Machine Learning Clusters on Campus
by: Xu, Kaiqiang, et al.
Published: (2021)
by: Xu, Kaiqiang, et al.
Published: (2021)
Joint Optimization of Training and Inference in Federated Edge Learning via Constrained Multi-Objective Deep Reinforcement Learning
by: Li, Zhen, et al.
Published: (2026)
by: Li, Zhen, et al.
Published: (2026)
Surviving the Edge: Federated Learning under Networking and Resource Constraints
by: Mwanje, Mike, et al.
Published: (2026)
by: Mwanje, Mike, et al.
Published: (2026)
Federated Learning and Evolutionary Game Model for Fog Federation Formation
by: Yasser, Zyad, et al.
Published: (2024)
by: Yasser, Zyad, et al.
Published: (2024)
Varuna: Enabling Failure-Type Aware RDMA Failover
by: Wang, Xiaoyang, et al.
Published: (2026)
by: Wang, Xiaoyang, et al.
Published: (2026)
Toward Co-adapting Machine Learning Job Shape and Cluster Topology
by: Chen, Shawn Shuoshuo, et al.
Published: (2025)
by: Chen, Shawn Shuoshuo, et al.
Published: (2025)
POSMAC: Powering Up In-Network AR/CG Traffic Classification with Online Learning
by: Shirmarz, Alireza, et al.
Published: (2025)
by: Shirmarz, Alireza, et al.
Published: (2025)
A Hybrid Cloud Management Plane for Data Processing Pipelines
by: Babu, Vignesh, et al.
Published: (2025)
by: Babu, Vignesh, et al.
Published: (2025)
The AutoSPADA Platform: User-Friendly Edge Computing for Distributed Learning and Data Analytics in Connected Vehicles
by: Nilsson, Adrian, et al.
Published: (2023)
by: Nilsson, Adrian, et al.
Published: (2023)
Similar Items
-
OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the Cloud
by: Warraich, Ertza, et al.
Published: (2023) -
Short-circuiting Rings for Low-Latency AllReduce
by: Hammer, Sarah-Michelle, et al.
Published: (2025) -
Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
by: Juerss, Anton, et al.
Published: (2026) -
Don't Let a Few Network Failures Slow the Entire AllReduce
by: Chen, Peiqing, et al.
Published: (2026) -
Rina: Enhancing Ring-AllReduce with In-network Aggregation in Distributed Model Training
by: Chen, Zixuan, et al.
Published: (2024)