Study of Workload Interference with Intelligent Routing on Dragonfly
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kang, Yao, Wang, Xin, Lan, Zhiling |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Q-adaptive: A Multi-Agent Reinforcement Learning Based Routing on Dragonfly Network
von: Kang, Yao, et al.
Veröffentlicht: (2024)
von: Kang, Yao, et al.
Veröffentlicht: (2024)
Leveraging InfiniBand Controller to Configure Deadlock-Free Routing Engines for Dragonflies
von: Maglione-Mathey, German, et al.
Veröffentlicht: (2025)
von: Maglione-Mathey, German, et al.
Veröffentlicht: (2025)
TINC: Trusted Intelligent NetChain
von: Xia, Qi, et al.
Veröffentlicht: (2025)
von: Xia, Qi, et al.
Veröffentlicht: (2025)
OptiNIC: A Resilient and Tail-Optimal RDMA NIC for Distributed ML Workloads
von: Warraich, Ertza, et al.
Veröffentlicht: (2025)
von: Warraich, Ertza, et al.
Veröffentlicht: (2025)
SAKURAONE: An Open Ethernet-Based AI HPC System and Its Observed Workload Dynamics in a Single-Tenant LLM Development Environment
von: Konishi, Fumikazu, et al.
Veröffentlicht: (2026)
von: Konishi, Fumikazu, et al.
Veröffentlicht: (2026)
ncsim: A Lightweight Simulator for Networked Edge Computing with Wireless Interference Modeling
von: Krishnamachari, Bhaskar, et al.
Veröffentlicht: (2026)
von: Krishnamachari, Bhaskar, et al.
Veröffentlicht: (2026)
An Open-Source Fast Parallel Routing Approach for Commercial FPGAs
von: Zang, Xinshi, et al.
Veröffentlicht: (2024)
von: Zang, Xinshi, et al.
Veröffentlicht: (2024)
A Lightweight and Scalable Design of Segment Routing in Broadband LEO Constellations Using Landmark-Based Skeleton Graphs
von: Hu, Menglan, et al.
Veröffentlicht: (2024)
von: Hu, Menglan, et al.
Veröffentlicht: (2024)
RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving
von: Kasnavieh, Hossein Hosseini, et al.
Veröffentlicht: (2026)
von: Kasnavieh, Hossein Hosseini, et al.
Veröffentlicht: (2026)
Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
von: Luo, Haoxiang, et al.
Veröffentlicht: (2025)
von: Luo, Haoxiang, et al.
Veröffentlicht: (2025)
Existence of Deadlock-Free Routing for Arbitrary Networks
von: Mendlovic, Uri, et al.
Veröffentlicht: (2025)
von: Mendlovic, Uri, et al.
Veröffentlicht: (2025)
FODT: Fast, Online, Distributed and Temporary Failure Recovery Approach for MEC
von: Yuan, Xin, et al.
Veröffentlicht: (2023)
von: Yuan, Xin, et al.
Veröffentlicht: (2023)
Varuna: Enabling Failure-Type Aware RDMA Failover
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2026)
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2026)
Collective Communication Profiling of Modern-day Machine Learning Workloads
von: Gupta, Jit, et al.
Veröffentlicht: (2025)
von: Gupta, Jit, et al.
Veröffentlicht: (2025)
A Multi-Cloud Framework for Zero-Trust Workload Authentication
von: Deochake, Saurabh, et al.
Veröffentlicht: (2025)
von: Deochake, Saurabh, et al.
Veröffentlicht: (2025)
From Skew to Symmetry: Node-Interconnect Multi-Path Balancing with Execution-time Planning for Modern GPU Clusters
von: Yao, Jinghan, et al.
Veröffentlicht: (2026)
von: Yao, Jinghan, et al.
Veröffentlicht: (2026)
Joint Network Slicing, Routing, and In-Network Computing for Energy-Efficient 6G
von: Sasan, Zeinab, et al.
Veröffentlicht: (2024)
von: Sasan, Zeinab, et al.
Veröffentlicht: (2024)
Edge-assisted Parallel Uncertain Skyline Processing for Low-latency IoE Analysis
von: Lai, Chuan-Chi, et al.
Veröffentlicht: (2025)
von: Lai, Chuan-Chi, et al.
Veröffentlicht: (2025)
OrchestrRL: Dynamic Compute and Network Orchestration for Disaggregated RL
von: Tan, Xin, et al.
Veröffentlicht: (2026)
von: Tan, Xin, et al.
Veröffentlicht: (2026)
Dynamic DAG-Application Scheduling for Multi-Tier Edge Computing in Heterogeneous Networks
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
A Study on 5G Network Slice Isolation Based on Native Cloud and Edge Computing Tools
von: Andrade, Maiko, et al.
Veröffentlicht: (2025)
von: Andrade, Maiko, et al.
Veröffentlicht: (2025)
Real-Time Scheduling for 802.1Qbv Time-Sensitive Networking (TSN): A Systematic Review and Experimental Study
von: Xue, Chuanyu, et al.
Veröffentlicht: (2023)
von: Xue, Chuanyu, et al.
Veröffentlicht: (2023)
D-LoRa: a Distributed Parameter Adaptation Scheme for LoRa Network
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
PSMOA: Policy Support Multi-Objective Optimization Algorithm for Decentralized Data Replication
von: Wang, Xi, et al.
Veröffentlicht: (2025)
von: Wang, Xi, et al.
Veröffentlicht: (2025)
Design and Operation of Shared Machine Learning Clusters on Campus
von: Xu, Kaiqiang, et al.
Veröffentlicht: (2021)
von: Xu, Kaiqiang, et al.
Veröffentlicht: (2021)
Multi-Source Coflow Scheduling in Collaborative Edge Computing with Multihop Network
von: Sahni, Yuvraj, et al.
Veröffentlicht: (2024)
von: Sahni, Yuvraj, et al.
Veröffentlicht: (2024)
Priority based inter-twin communication in vehicular digital twin networks
von: Zia, Qasim, et al.
Veröffentlicht: (2024)
von: Zia, Qasim, et al.
Veröffentlicht: (2024)
Fast Multichannel Topology Discovery in Cognitive Radio Networks
von: Wang, Yung-Li, et al.
Veröffentlicht: (2025)
von: Wang, Yung-Li, et al.
Veröffentlicht: (2025)
Bodyless Block Propagation: TPS Fully Scalable Blockchain with Pre-Validation
von: Zhao, Chonghe, et al.
Veröffentlicht: (2022)
von: Zhao, Chonghe, et al.
Veröffentlicht: (2022)
BlockSDN: Towards a High-Performance Blockchain via Software-Defined Cross Networking optimization
von: Jia, Wenyang, et al.
Veröffentlicht: (2025)
von: Jia, Wenyang, et al.
Veröffentlicht: (2025)
Towards Timely Video Analytics Services at the Network Edge
von: Li, Xishuo, et al.
Veröffentlicht: (2024)
von: Li, Xishuo, et al.
Veröffentlicht: (2024)
EdgeTimer: Adaptive Multi-Timescale Scheduling in Mobile Edge Computing with Deep Reinforcement Learning
von: Hao, Yijun, et al.
Veröffentlicht: (2024)
von: Hao, Yijun, et al.
Veröffentlicht: (2024)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
von: Du, Chengze, et al.
Veröffentlicht: (2025)
von: Du, Chengze, et al.
Veröffentlicht: (2025)
Accelerating Stable Matching between Workers and Spatial-Temporal Tasks for Dynamic MCS: A Stagewise Service Trading Approach
von: Qi, Houyi, et al.
Veröffentlicht: (2025)
von: Qi, Houyi, et al.
Veröffentlicht: (2025)
Diving into 3D Parallelism with Heterogeneous Spot Instance GPUs: Design and Implications
von: Wang, Yuxiao, et al.
Veröffentlicht: (2025)
von: Wang, Yuxiao, et al.
Veröffentlicht: (2025)
Towards Practical Overlay Networks for Decentralized Federated Learning
von: Hua, Yifan, et al.
Veröffentlicht: (2024)
von: Hua, Yifan, et al.
Veröffentlicht: (2024)
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
von: Xu, Heng, et al.
Veröffentlicht: (2025)
von: Xu, Heng, et al.
Veröffentlicht: (2025)
PeerSync: Accelerating Containerized Service Delivery at the Network Edge
von: Deng, Yinuo, et al.
Veröffentlicht: (2025)
von: Deng, Yinuo, et al.
Veröffentlicht: (2025)
Recursive Offloading for LLM Serving in Multi-tier Networks
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
Arcturus: A Cloud Overlay Network for Global Accelerator with Enhanced Performance and Stability
von: Liu, Matthew Yang, et al.
Veröffentlicht: (2025)
von: Liu, Matthew Yang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Q-adaptive: A Multi-Agent Reinforcement Learning Based Routing on Dragonfly Network
von: Kang, Yao, et al.
Veröffentlicht: (2024) -
Leveraging InfiniBand Controller to Configure Deadlock-Free Routing Engines for Dragonflies
von: Maglione-Mathey, German, et al.
Veröffentlicht: (2025) -
TINC: Trusted Intelligent NetChain
von: Xia, Qi, et al.
Veröffentlicht: (2025) -
OptiNIC: A Resilient and Tail-Optimal RDMA NIC for Distributed ML Workloads
von: Warraich, Ertza, et al.
Veröffentlicht: (2025) -
SAKURAONE: An Open Ethernet-Based AI HPC System and Its Observed Workload Dynamics in a Single-Tenant LLM Development Environment
von: Konishi, Fumikazu, et al.
Veröffentlicht: (2026)