Cluster Topology-Driven Placement of Experts Reduces Network Traffic in MoE Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Sivtsov, Danil, Katrutsa, Aleksandr, Oseledets, Ivan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpaceMoE: Realizing Distributed Mixture-of-Experts Inference over Space Networks
by: Wang, Zhanwei, et al.
Published: (2026)
by: Wang, Zhanwei, et al.
Published: (2026)
LOAM: Low-latency Communication, Caching, and Computation Placement in Data-Intensive Computing Networks
by: Zhang, Jinkun, et al.
Published: (2024)
by: Zhang, Jinkun, et al.
Published: (2024)
Toward Co-adapting Machine Learning Job Shape and Cluster Topology
by: Chen, Shawn Shuoshuo, et al.
Published: (2025)
by: Chen, Shawn Shuoshuo, et al.
Published: (2025)
Optimal Oblivious Load-Balancing for Sparse Traffic in Large-Scale Satellite Networks
by: Ramakanth, Rudrapatna Vallabh, et al.
Published: (2026)
by: Ramakanth, Rudrapatna Vallabh, et al.
Published: (2026)
POSMAC: Powering Up In-Network AR/CG Traffic Classification with Online Learning
by: Shirmarz, Alireza, et al.
Published: (2025)
by: Shirmarz, Alireza, et al.
Published: (2025)
Revolutionizing Datacenter Networks via Reconfigurable Topologies
by: Avin, Chen, et al.
Published: (2025)
by: Avin, Chen, et al.
Published: (2025)
Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
by: Juerss, Anton, et al.
Published: (2026)
by: Juerss, Anton, et al.
Published: (2026)
Fast Multichannel Topology Discovery in Cognitive Radio Networks
by: Wang, Yung-Li, et al.
Published: (2025)
by: Wang, Yung-Li, et al.
Published: (2025)
A New Broadcast Model for Several Network Topologies
by: Lu, Hongbo, et al.
Published: (2025)
by: Lu, Hongbo, et al.
Published: (2025)
HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge Network
by: Zheng, Peirong, et al.
Published: (2026)
by: Zheng, Peirong, et al.
Published: (2026)
Legible Consensus: Topology-Aware Quorum Geometry for Asymmetric Networks
by: Mason, Tony
Published: (2026)
by: Mason, Tony
Published: (2026)
Efficient Fog Node Placement using Nature-Inspired Metaheuristic for IoT Applications
by: Naouri, Abdenacer, et al.
Published: (2023)
by: Naouri, Abdenacer, et al.
Published: (2023)
Decentralized Network Topology Design for Task Offloading in Mobile Edge Computing
by: Ma, Ke, et al.
Published: (2024)
by: Ma, Ke, et al.
Published: (2024)
Topology-aware Microservice Architecture in Edge Networks: Deployment Optimization and Implementation
by: Chen, Yuang, et al.
Published: (2025)
by: Chen, Yuang, et al.
Published: (2025)
Resource Allocation Driven by Large Models in Future Semantic-Aware Networks
by: Zhang, Haijun, et al.
Published: (2025)
by: Zhang, Haijun, et al.
Published: (2025)
OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the Cloud
by: Warraich, Ertza, et al.
Published: (2023)
by: Warraich, Ertza, et al.
Published: (2023)
A New Classification of Clustering-based for Different Problems in Different Wireless Ad-hoc Networks
by: Boualem, Adda, et al.
Published: (2024)
by: Boualem, Adda, et al.
Published: (2024)
FlowTracer: A Tool for Uncovering Network Path Usage Imbalance in AI Training Clusters
by: Jamil, Hasibul, et al.
Published: (2024)
by: Jamil, Hasibul, et al.
Published: (2024)
SplitLLM: Collaborative Inference of LLMs for Model Placement and Throughput Optimization
by: Mudvari, Akrit, et al.
Published: (2024)
by: Mudvari, Akrit, et al.
Published: (2024)
A Task Decomposition and Planning Framework for Efficient LLM Inference in AI-Enabled WiFi-Offload Networks
by: Han, Mingqi, et al.
Published: (2026)
by: Han, Mingqi, et al.
Published: (2026)
Rina: Enhancing Ring-AllReduce with In-network Aggregation in Distributed Model Training
by: Chen, Zixuan, et al.
Published: (2024)
by: Chen, Zixuan, et al.
Published: (2024)
AllReduce Scheduling with Hierarchical Deep Reinforcement Learning
by: Wei, Yufan, et al.
Published: (2025)
by: Wei, Yufan, et al.
Published: (2025)
Short-circuiting Rings for Low-Latency AllReduce
by: Hammer, Sarah-Michelle, et al.
Published: (2025)
by: Hammer, Sarah-Michelle, et al.
Published: (2025)
CRAFT: Latency and Cost-Aware Genetic-Based Framework for Node Placement in Edge-Fog Environments
by: Mahdizadeh, Soheil, et al.
Published: (2025)
by: Mahdizadeh, Soheil, et al.
Published: (2025)
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
by: Xu, Heng, et al.
Published: (2025)
by: Xu, Heng, et al.
Published: (2025)
Topological Analysis for Identifying Anomalies in Serverless Platforms
by: Reali, Gianluca, et al.
Published: (2026)
by: Reali, Gianluca, et al.
Published: (2026)
TraDE: Network and Traffic-aware Adaptive Scheduling for Microservices Under Dynamics
by: Chen, Ming, et al.
Published: (2024)
by: Chen, Ming, et al.
Published: (2024)
Trust-Aware Routing for Distributed Generative AI Inference at the Edge
by: Nguyen, Chanh, et al.
Published: (2026)
by: Nguyen, Chanh, et al.
Published: (2026)
Optimizing Resource Allocation for Geographically-Distributed Inference by Large Language Models
by: Sun, Tingyang, et al.
Published: (2025)
by: Sun, Tingyang, et al.
Published: (2025)
Design and Operation of Shared Machine Learning Clusters on Campus
by: Xu, Kaiqiang, et al.
Published: (2021)
by: Xu, Kaiqiang, et al.
Published: (2021)
AI Greenferencing: Routing AI Inferencing to Green Modular Data Centers with Heron
by: Reddy, Tella Rajashekhar, et al.
Published: (2025)
by: Reddy, Tella Rajashekhar, et al.
Published: (2025)
Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
by: Ye, Shengyuan, et al.
Published: (2025)
by: Ye, Shengyuan, et al.
Published: (2025)
Efficient All-to-All Collective Communication Schedules for Direct-Connect Topologies
by: Basu, Prithwish, et al.
Published: (2023)
by: Basu, Prithwish, et al.
Published: (2023)
YUHENG-OS: A Cloud-Native Space Cluster Operating System
by: Zhang, Jin, et al.
Published: (2026)
by: Zhang, Jin, et al.
Published: (2026)
XWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy Farms
by: Reddy, Tella Rajashekhar, et al.
Published: (2026)
by: Reddy, Tella Rajashekhar, et al.
Published: (2026)
ClusterSlice: A Zero-touch Deployment Platform for the Edge Cloud Continuum
by: Mamatas, Lefteris, et al.
Published: (2024)
by: Mamatas, Lefteris, et al.
Published: (2024)
SlimCaching: Edge Caching of Mixture-of-Experts for Distributed Inference
by: Chen, Qian, et al.
Published: (2025)
by: Chen, Qian, et al.
Published: (2025)
Causal Inference for Quantifying Noisy Neighbor Effects in Multi-Tenant Cloud Environments
by: Schiavo, Philipe S., et al.
Published: (2026)
by: Schiavo, Philipe S., et al.
Published: (2026)
LIDC: A Location Independent Multi-Cluster Computing Framework for Data Intensive Science
by: Timilsina, Sankalpa, et al.
Published: (2025)
by: Timilsina, Sankalpa, et al.
Published: (2025)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
by: Du, Chengze, et al.
Published: (2025)
by: Du, Chengze, et al.
Published: (2025)
Similar Items
-
SpaceMoE: Realizing Distributed Mixture-of-Experts Inference over Space Networks
by: Wang, Zhanwei, et al.
Published: (2026) -
LOAM: Low-latency Communication, Caching, and Computation Placement in Data-Intensive Computing Networks
by: Zhang, Jinkun, et al.
Published: (2024) -
Toward Co-adapting Machine Learning Job Shape and Cluster Topology
by: Chen, Shawn Shuoshuo, et al.
Published: (2025) -
Optimal Oblivious Load-Balancing for Sparse Traffic in Large-Scale Satellite Networks
by: Ramakanth, Rudrapatna Vallabh, et al.
Published: (2026) -
POSMAC: Powering Up In-Network AR/CG Traffic Classification with Online Learning
by: Shirmarz, Alireza, et al.
Published: (2025)