ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Wenxiang, Pan, Xinglin, Fan, Ruibo, Shi, Shaohuai, Chu, Xiaowen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
von: Lin, Wenxiang, et al.
Veröffentlicht: (2025)
von: Lin, Wenxiang, et al.
Veröffentlicht: (2025)
Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules
von: Pan, Xinglin, et al.
Veröffentlicht: (2024)
von: Pan, Xinglin, et al.
Veröffentlicht: (2024)
DreamDDP: Accelerating Data Parallel Distributed LLM Training with Layer-wise Scheduled Partial Synchronization
von: Tang, Zhenheng, et al.
Veröffentlicht: (2025)
von: Tang, Zhenheng, et al.
Veröffentlicht: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
UCCL-Zip: Lossless Compression Supercharged GPU Communication
von: Ma, Shuang, et al.
Veröffentlicht: (2026)
von: Ma, Shuang, et al.
Veröffentlicht: (2026)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
Bandwidth-Aware and Overlap-Weighted Compression for Communication-Efficient Federated Learning
von: Tang, Zichen, et al.
Veröffentlicht: (2024)
von: Tang, Zichen, et al.
Veröffentlicht: (2024)
ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and Compression
von: Wang, Zirui, et al.
Veröffentlicht: (2025)
von: Wang, Zirui, et al.
Veröffentlicht: (2025)
FusionLLM: A Decentralized LLM Training System on Geo-distributed GPUs with Adaptive Compression
von: Tang, Zhenheng, et al.
Veröffentlicht: (2024)
von: Tang, Zhenheng, et al.
Veröffentlicht: (2024)
HiCCL: A Hierarchical Collective Communication Library
von: Hidayetoglu, Mert, et al.
Veröffentlicht: (2024)
von: Hidayetoglu, Mert, et al.
Veröffentlicht: (2024)
AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs
von: Lin, Wenxiang, et al.
Veröffentlicht: (2026)
von: Lin, Wenxiang, et al.
Veröffentlicht: (2026)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
von: Kim, Heehoon, et al.
Veröffentlicht: (2026)
von: Kim, Heehoon, et al.
Veröffentlicht: (2026)
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
von: Guo, Yipin, et al.
Veröffentlicht: (2026)
von: Guo, Yipin, et al.
Veröffentlicht: (2026)
gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters
von: Huang, Jiajun, et al.
Veröffentlicht: (2023)
von: Huang, Jiajun, et al.
Veröffentlicht: (2023)
Floating-Point Data Transformation for Lossless Compression
von: Jamalidinan, Samirasadat, et al.
Veröffentlicht: (2025)
von: Jamalidinan, Samirasadat, et al.
Veröffentlicht: (2025)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
von: Sun, Mingyu, et al.
Veröffentlicht: (2025)
von: Sun, Mingyu, et al.
Veröffentlicht: (2025)
ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling
von: Yang, Yuchen, et al.
Veröffentlicht: (2026)
von: Yang, Yuchen, et al.
Veröffentlicht: (2026)
Boosting Scientific Error-Bounded Lossy Compression through Optimized Synergistic Lossy-Lossless Orchestration
von: Wu, Shixun, et al.
Veröffentlicht: (2025)
von: Wu, Shixun, et al.
Veröffentlicht: (2025)
Exploiting Multicast for Accelerating Collective Communication
von: Xu, Chao, et al.
Veröffentlicht: (2026)
von: Xu, Chao, et al.
Veröffentlicht: (2026)
Accelerating Compound LLM Training Workloads with Maestro
von: Yuan, Xiulong, et al.
Veröffentlicht: (2026)
von: Yuan, Xiulong, et al.
Veröffentlicht: (2026)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
von: Zhang, Mingjun, et al.
Veröffentlicht: (2025)
von: Zhang, Mingjun, et al.
Veröffentlicht: (2025)
AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
von: Chen, Qiaoling, et al.
Veröffentlicht: (2023)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2023)
ZCCL: Significantly Improving Collective Communication With Error-Bounded Lossy Compression
von: Huang, Jiajun, et al.
Veröffentlicht: (2025)
von: Huang, Jiajun, et al.
Veröffentlicht: (2025)
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
von: Liu, Man, et al.
Veröffentlicht: (2026)
von: Liu, Man, et al.
Veröffentlicht: (2026)
Accelerating Communication in Deep Learning Recommendation Model Training with Dual-Level Adaptive Lossy Compression
von: Feng, Hao, et al.
Veröffentlicht: (2024)
von: Feng, Hao, et al.
Veröffentlicht: (2024)
PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep Learning
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
von: Wang, Yisu, et al.
Veröffentlicht: (2025)
FourierCompress: Layer-Aware Spectral Activation Compression for Efficient and Accurate Collaborative LLM Inference
von: Ma, Jian, et al.
Veröffentlicht: (2025)
von: Ma, Jian, et al.
Veröffentlicht: (2025)
70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2025)
Scheduling Deep Learning Jobs in Multi-Tenant GPU Clusters via Wise Resource Sharing
von: Luo, Yizhou, et al.
Veröffentlicht: (2024)
von: Luo, Yizhou, et al.
Veröffentlicht: (2024)
PipeDiT: Accelerating Diffusion Transformers in Video Generation with Task Pipelining and Model Decoupling
von: Wang, Sijie, et al.
Veröffentlicht: (2025)
von: Wang, Sijie, et al.
Veröffentlicht: (2025)
FedImpro: Measuring and Improving Client Update in Federated Learning
von: Tang, Zhenheng, et al.
Veröffentlicht: (2024)
von: Tang, Zhenheng, et al.
Veröffentlicht: (2024)
AcceLLM: Accelerating LLM Inference using Redundancy for Load Balancing and Data Locality
von: Bournias, Ilias, et al.
Veröffentlicht: (2024)
von: Bournias, Ilias, et al.
Veröffentlicht: (2024)
ZipFlow: a Compiler-based Framework to Unleash Compressed Data Movement for Modern GPUs
von: Yeo, Gwangoo, et al.
Veröffentlicht: (2026)
von: Yeo, Gwangoo, et al.
Veröffentlicht: (2026)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs
von: Yao, Xiaozhe, et al.
Veröffentlicht: (2023)
von: Yao, Xiaozhe, et al.
Veröffentlicht: (2023)
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training
von: Deng, Yangtao, et al.
Veröffentlicht: (2025)
von: Deng, Yangtao, et al.
Veröffentlicht: (2025)
PALM: A Efficient Performance Simulator for Tiled Accelerators with Large-scale Model Training
von: Fang, Jiahao, et al.
Veröffentlicht: (2024)
von: Fang, Jiahao, et al.
Veröffentlicht: (2024)
Pier: Efficient Large Language Model pretraining with Relaxed Global Communication
von: Fan, Shuyuan, et al.
Veröffentlicht: (2025)
von: Fan, Shuyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
von: Fan, Ruibo, et al.
Veröffentlicht: (2026) -
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
von: Lin, Wenxiang, et al.
Veröffentlicht: (2025) -
Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules
von: Pan, Xinglin, et al.
Veröffentlicht: (2024) -
DreamDDP: Accelerating Data Parallel Distributed LLM Training with Layer-wise Scheduled Partial Synchronization
von: Tang, Zhenheng, et al.
Veröffentlicht: (2025) -
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)