Exploiting Multicast for Accelerating Collective Communication
Fuente:
arXiv
Guardado en:
| Autores principales: | Xu, Chao, Zhang, Xu, Luo, Zihang, Wu, Yuyan, Qian, Guoxin, Yao, Yufeng, Wang, Chihyung, Zhou, Jingbin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters
por: Huang, Jiajun, et al.
Publicado: (2023)
por: Huang, Jiajun, et al.
Publicado: (2023)
FUSCO: High-Performance Distributed Data Shuffling via Transformation-Communication Fusion
por: Zhu, Zhuoran, et al.
Publicado: (2025)
por: Zhu, Zhuoran, et al.
Publicado: (2025)
Rubick: Exploiting Job Reconfigurability for Deep Learning Cluster Scheduling
por: Zhang, Xinyi, et al.
Publicado: (2024)
por: Zhang, Xinyi, et al.
Publicado: (2024)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
por: Wang, Zhigang, et al.
Publicado: (2024)
por: Wang, Zhigang, et al.
Publicado: (2024)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
por: Zhang, Mingjun, et al.
Publicado: (2025)
por: Zhang, Mingjun, et al.
Publicado: (2025)
Air-FedGA: A Grouping Asynchronous Federated Learning Mechanism Exploiting Over-the-air Computation
por: Ma, Qianpiao, et al.
Publicado: (2025)
por: Ma, Qianpiao, et al.
Publicado: (2025)
Cross-region Model Training with Communication-Computation Overlapping and Delay Compensation
por: Zhu, Ying, et al.
Publicado: (2025)
por: Zhu, Ying, et al.
Publicado: (2025)
EPIC: Abstraction and Polymorphism of In-Network Collectives on Ethernet
por: Yuan, Yitao, et al.
Publicado: (2026)
por: Yuan, Yitao, et al.
Publicado: (2026)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
por: Wu, Tian, et al.
Publicado: (2025)
por: Wu, Tian, et al.
Publicado: (2025)
Distributed Consensus Network: A Modularized Communication Framework and Reliability Probabilistic Analysis
por: Li, Yuetai, et al.
Publicado: (2025)
por: Li, Yuetai, et al.
Publicado: (2025)
On-the-fly Communication-and-Computing to Enable Representation Learning for Distributed Point Clouds
por: Chen, Xu, et al.
Publicado: (2024)
por: Chen, Xu, et al.
Publicado: (2024)
ZCCL: Significantly Improving Collective Communication With Error-Bounded Lossy Compression
por: Huang, Jiajun, et al.
Publicado: (2025)
por: Huang, Jiajun, et al.
Publicado: (2025)
Accelerating Distributed MoE Training and Inference with Lina
por: Li, Jiamin, et al.
Publicado: (2022)
por: Li, Jiamin, et al.
Publicado: (2022)
Accelerating Compound LLM Training Workloads with Maestro
por: Yuan, Xiulong, et al.
Publicado: (2026)
por: Yuan, Xiulong, et al.
Publicado: (2026)
Prime Collective Communications Library -- Technical Report
por: Keiblinger, Michael, et al.
Publicado: (2025)
por: Keiblinger, Michael, et al.
Publicado: (2025)
AES-SpMM: Balancing Accuracy and Speed by Adaptive Edge Sampling Strategy to Accelerate SpMM in GNNs
por: Song, Yingchen, et al.
Publicado: (2025)
por: Song, Yingchen, et al.
Publicado: (2025)
GPU-Accelerated Distributed QAOA on Large-scale HPC Ecosystems
por: Xu, Zhihao, et al.
Publicado: (2025)
por: Xu, Zhihao, et al.
Publicado: (2025)
Hyperion: Hierarchical Scheduling for Parallel LLM Acceleration in Multi-tier Networks
por: Ma, Mulei, et al.
Publicado: (2025)
por: Ma, Mulei, et al.
Publicado: (2025)
HiCCL: A Hierarchical Collective Communication Library
por: Hidayetoglu, Mert, et al.
Publicado: (2024)
por: Hidayetoglu, Mert, et al.
Publicado: (2024)
Accelerating End-Cloud Collaborative Inference via Near Bubble-free Pipeline Optimization
por: Gao, Luyao, et al.
Publicado: (2024)
por: Gao, Luyao, et al.
Publicado: (2024)
ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training
por: Lin, Wenxiang, et al.
Publicado: (2026)
por: Lin, Wenxiang, et al.
Publicado: (2026)
Union: An Automatic Workload Manager for Accelerating Network Simulation
por: Wang, Xin, et al.
Publicado: (2024)
por: Wang, Xin, et al.
Publicado: (2024)
A HPX Communication Benchmark: Distributed FFT using Collectives
por: Strack, Alexander, et al.
Publicado: (2025)
por: Strack, Alexander, et al.
Publicado: (2025)
Collaborative Inference Acceleration with Non-Penetrative Tensor Partitioning
por: Liu, Zhibang, et al.
Publicado: (2025)
por: Liu, Zhibang, et al.
Publicado: (2025)
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training
por: Deng, Yangtao, et al.
Publicado: (2025)
por: Deng, Yangtao, et al.
Publicado: (2025)
Extracting the Potential of Emerging Hardware Accelerators for Symmetric Eigenvalue Decomposition
por: Wang, Hansheng, et al.
Publicado: (2024)
por: Wang, Hansheng, et al.
Publicado: (2024)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
por: Wang, Zhibin, et al.
Publicado: (2025)
por: Wang, Zhibin, et al.
Publicado: (2025)
exa-AMD: A Scalable Workflow for Accelerating AI-Assisted Materials Discovery and Design
por: Moraru, Maxim, et al.
Publicado: (2025)
por: Moraru, Maxim, et al.
Publicado: (2025)
A Portable Framework for Accelerating Stencil Computations on Modern Node Architectures
por: Sai, Ryuichi, et al.
Publicado: (2023)
por: Sai, Ryuichi, et al.
Publicado: (2023)
Accelerating OpenPangu Inference on NPU via Speculative Decoding
por: Dai, Yuntao, et al.
Publicado: (2026)
por: Dai, Yuntao, et al.
Publicado: (2026)
Generic Multicast (Extended Version)
por: Bolina, José Augusto, et al.
Publicado: (2024)
por: Bolina, José Augusto, et al.
Publicado: (2024)
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching
por: Zhao, Zhan, et al.
Publicado: (2026)
por: Zhao, Zhan, et al.
Publicado: (2026)
Exploiting Stragglers in Distributed Computing Systems with Task Grouping
por: Adikari, Tharindu, et al.
Publicado: (2024)
por: Adikari, Tharindu, et al.
Publicado: (2024)
PALM: A Efficient Performance Simulator for Tiled Accelerators with Large-scale Model Training
por: Fang, Jiahao, et al.
Publicado: (2024)
por: Fang, Jiahao, et al.
Publicado: (2024)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
por: Chen, Liangkun, et al.
Publicado: (2025)
por: Chen, Liangkun, et al.
Publicado: (2025)
OmniInfer: System-Wide Acceleration Techniques for Optimizing LLM Serving Throughput and Latency
por: Wang, Jun, et al.
Publicado: (2025)
por: Wang, Jun, et al.
Publicado: (2025)
Efficient Local-to-Global Collaborative Perception via Joint Communication and Computation Optimization
por: Zhang, Hui, et al.
Publicado: (2026)
por: Zhang, Hui, et al.
Publicado: (2026)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
por: Xu, Guanbin, et al.
Publicado: (2026)
por: Xu, Guanbin, et al.
Publicado: (2026)
Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution
por: Wang, Siqi, et al.
Publicado: (2024)
por: Wang, Siqi, et al.
Publicado: (2024)
Collaborative Inference for Large Models with Task Offloading and Early Exiting
por: Xie, Zuan, et al.
Publicado: (2024)
por: Xie, Zuan, et al.
Publicado: (2024)
Ejemplares similares
-
gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters
por: Huang, Jiajun, et al.
Publicado: (2023) -
FUSCO: High-Performance Distributed Data Shuffling via Transformation-Communication Fusion
por: Zhu, Zhuoran, et al.
Publicado: (2025) -
Rubick: Exploiting Job Reconfigurability for Deep Learning Cluster Scheduling
por: Zhang, Xinyi, et al.
Publicado: (2024) -
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
por: Wang, Zhigang, et al.
Publicado: (2024) -
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
por: Zhang, Mingjun, et al.
Publicado: (2025)