Reliable and Resilient Collective Communication Library for LLM Training and Serving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Wei, Yu, Nengneng, Xiong, Sixian, Liu, Zaoxing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Don't Let a Few Network Failures Slow the Entire AllReduce
von: Chen, Peiqing, et al.
Veröffentlicht: (2026)
von: Chen, Peiqing, et al.
Veröffentlicht: (2026)
Efficient Direct-Connect Topologies for Collective Communications
von: Zhao, Liangyu, et al.
Veröffentlicht: (2022)
von: Zhao, Liangyu, et al.
Veröffentlicht: (2022)
ForestColl: Throughput-Optimal Collective Communications on Heterogeneous Network Fabrics
von: Zhao, Liangyu, et al.
Veröffentlicht: (2024)
von: Zhao, Liangyu, et al.
Veröffentlicht: (2024)
Meili: Enabling SmartNIC as a Service in the Cloud
von: Su, Qiang, et al.
Veröffentlicht: (2023)
von: Su, Qiang, et al.
Veröffentlicht: (2023)
Recursive Offloading for LLM Serving in Multi-tier Networks
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
FedRDMA: Communication-Efficient Cross-Silo Federated LLM via Chunked RDMA Transmission
von: Zhang, Zeling, et al.
Veröffentlicht: (2024)
von: Zhang, Zeling, et al.
Veröffentlicht: (2024)
Architectural Blueprint For Heterogeneity-Resilient Federated Learning
von: Bashir, Satwat, et al.
Veröffentlicht: (2024)
von: Bashir, Satwat, et al.
Veröffentlicht: (2024)
Multi-stage Flow Scheduling for LLM Serving
von: Sun, Yijun, et al.
Veröffentlicht: (2026)
von: Sun, Yijun, et al.
Veröffentlicht: (2026)
InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching Transceivers
von: Shou, Chenchen, et al.
Veröffentlicht: (2025)
von: Shou, Chenchen, et al.
Veröffentlicht: (2025)
MLTCP: Congestion Control for DNN Training
von: Rajasekaran, Sudarsanan, et al.
Veröffentlicht: (2024)
von: Rajasekaran, Sudarsanan, et al.
Veröffentlicht: (2024)
SFL-LEO: Asynchronous Split-Federated Learning Design for LEO Satellite-Ground Network Framework
von: Wu, Jiasheng, et al.
Veröffentlicht: (2025)
von: Wu, Jiasheng, et al.
Veröffentlicht: (2025)
RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving
von: Kasnavieh, Hossein Hosseini, et al.
Veröffentlicht: (2026)
von: Kasnavieh, Hossein Hosseini, et al.
Veröffentlicht: (2026)
SplitLLM: Collaborative Inference of LLMs for Model Placement and Throughput Optimization
von: Mudvari, Akrit, et al.
Veröffentlicht: (2024)
von: Mudvari, Akrit, et al.
Veröffentlicht: (2024)
FedMFS: Federated Multimodal Fusion Learning with Selective Modality Communication
von: Yuan, Liangqi, et al.
Veröffentlicht: (2023)
von: Yuan, Liangqi, et al.
Veröffentlicht: (2023)
KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
von: Liu, Zedong, et al.
Veröffentlicht: (2026)
von: Liu, Zedong, et al.
Veröffentlicht: (2026)
Reactive Orchestration for Hierarchical Federated Learning Under a Communication Cost Budget
von: Čilić, Ivan, et al.
Veröffentlicht: (2024)
von: Čilić, Ivan, et al.
Veröffentlicht: (2024)
COMSPLIT: A Communication-Aware Split Learning Design for Heterogeneous IoT Platforms
von: Ninkovic, Vukan, et al.
Veröffentlicht: (2024)
von: Ninkovic, Vukan, et al.
Veröffentlicht: (2024)
Joint Optimization of Training and Inference in Federated Edge Learning via Constrained Multi-Objective Deep Reinforcement Learning
von: Li, Zhen, et al.
Veröffentlicht: (2026)
von: Li, Zhen, et al.
Veröffentlicht: (2026)
FedSkipTwin: Digital-Twin-Guided Client Skipping for Communication-Efficient Federated Learning
von: Commey, Daniel, et al.
Veröffentlicht: (2025)
von: Commey, Daniel, et al.
Veröffentlicht: (2025)
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
von: Xu, Heng, et al.
Veröffentlicht: (2025)
von: Xu, Heng, et al.
Veröffentlicht: (2025)
Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
von: Luo, Haoxiang, et al.
Veröffentlicht: (2025)
von: Luo, Haoxiang, et al.
Veröffentlicht: (2025)
Efficient All-to-All Collective Communication Schedules for Direct-Connect Topologies
von: Basu, Prithwish, et al.
Veröffentlicht: (2023)
von: Basu, Prithwish, et al.
Veröffentlicht: (2023)
Network Digital Untwinning: Towards Backward Optimization of Digital Twins
von: Zhang, Zifan, et al.
Veröffentlicht: (2026)
von: Zhang, Zifan, et al.
Veröffentlicht: (2026)
Harvest: Adaptive Photonic Switching Schedules for Collective Communication in Scale-up Domains
von: Rahman, Mahir, et al.
Veröffentlicht: (2026)
von: Rahman, Mahir, et al.
Veröffentlicht: (2026)
Hierarchical Online-Scheduling for Energy-Efficient Split Inference with Progressive Transmission
von: Tang, Zengzipeng, et al.
Veröffentlicht: (2026)
von: Tang, Zengzipeng, et al.
Veröffentlicht: (2026)
EMA: Efficient Model Adaptation for Learning-based Systems
von: Yu, Daiyang, et al.
Veröffentlicht: (2026)
von: Yu, Daiyang, et al.
Veröffentlicht: (2026)
Federated Split Learning with Model Pruning and Gradient Quantization in Wireless Networks
von: Zhang, Junhe, et al.
Veröffentlicht: (2024)
von: Zhang, Junhe, et al.
Veröffentlicht: (2024)
RailX: A Flexible, Scalable, and Low-Cost Network Architecture for Hyper-Scale LLM Training Systems
von: Feng, Yinxiao, et al.
Veröffentlicht: (2025)
von: Feng, Yinxiao, et al.
Veröffentlicht: (2025)
Optimizing Edge Offloading Decisions for Object Detection
von: Qiu, Jiaming, et al.
Veröffentlicht: (2024)
von: Qiu, Jiaming, et al.
Veröffentlicht: (2024)
Device Sampling and Resource Optimization for Federated Learning in Cooperative Edge Networks
von: Wang, Su, et al.
Veröffentlicht: (2023)
von: Wang, Su, et al.
Veröffentlicht: (2023)
MergeSFL: Split Federated Learning with Feature Merging and Batch Size Regulation
von: Liao, Yunming, et al.
Veröffentlicht: (2023)
von: Liao, Yunming, et al.
Veröffentlicht: (2023)
Satellite Federated Fine-Tuning for Foundation Models in Space Computing Power Networks
von: Zhu, Yan, et al.
Veröffentlicht: (2025)
von: Zhu, Yan, et al.
Veröffentlicht: (2025)
The Internet of Things in the Era of Generative AI: Vision and Challenges
von: Wang, Xin, et al.
Veröffentlicht: (2024)
von: Wang, Xin, et al.
Veröffentlicht: (2024)
TinyLLM: A Framework for Training and Deploying Language Models at the Edge Computers
von: Kandala, Savitha Viswanadh, et al.
Veröffentlicht: (2024)
von: Kandala, Savitha Viswanadh, et al.
Veröffentlicht: (2024)
Distributed Simulation for Digital Twins of Large-Scale Real-World DiffServ-Based Networks
von: Huang, Zhuoyao, et al.
Veröffentlicht: (2024)
von: Huang, Zhuoyao, et al.
Veröffentlicht: (2024)
DRL-Based Federated Self-Supervised Learning for Task Offloading and Resource Allocation in ISAC-Enabled Vehicle Edge Computing
von: Gu, Xueying, et al.
Veröffentlicht: (2024)
von: Gu, Xueying, et al.
Veröffentlicht: (2024)
ZKP-FedEval: Verifiable and Privacy-Preserving Federated Evaluation using Zero-Knowledge Proofs
von: Commey, Daniel, et al.
Veröffentlicht: (2025)
von: Commey, Daniel, et al.
Veröffentlicht: (2025)
Online Identification of IT Systems through Active Causal Learning
von: Hammar, Kim, et al.
Veröffentlicht: (2025)
von: Hammar, Kim, et al.
Veröffentlicht: (2025)
Flooding with Absorption: An Efficient Protocol for Heterogeneous Bandits over Complex Networks
von: Lee, Junghyun, et al.
Veröffentlicht: (2023)
von: Lee, Junghyun, et al.
Veröffentlicht: (2023)
Improving Slow Transfer Predictions: Generative Methods Compared
von: Kim, Jacob Taegon, et al.
Veröffentlicht: (2025)
von: Kim, Jacob Taegon, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Don't Let a Few Network Failures Slow the Entire AllReduce
von: Chen, Peiqing, et al.
Veröffentlicht: (2026) -
Efficient Direct-Connect Topologies for Collective Communications
von: Zhao, Liangyu, et al.
Veröffentlicht: (2022) -
ForestColl: Throughput-Optimal Collective Communications on Heterogeneous Network Fabrics
von: Zhao, Liangyu, et al.
Veröffentlicht: (2024) -
Meili: Enabling SmartNIC as a Service in the Cloud
von: Su, Qiang, et al.
Veröffentlicht: (2023) -
Recursive Offloading for LLM Serving in Multi-tier Networks
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)