UCCL-Zip: Lossless Compression Supercharged GPU Communication
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Shuang, Lao, Chon Lam, Xu, Zhiying, Wang, Zhuang, Mao, Ziming, Meng, Delong, Zhen, Jia, Wu, Jun, Stoica, Ion, Wang, Yida, Zhou, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UCCL-EP: Portable Expert-Parallel Communication
by: Mao, Ziming, et al.
Published: (2025)
by: Mao, Ziming, et al.
Published: (2025)
ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training
by: Lin, Wenxiang, et al.
Published: (2026)
by: Lin, Wenxiang, et al.
Published: (2026)
Pie: Pooling CPU Memory for LLM Inference
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
by: Guo, Yipin, et al.
Published: (2026)
by: Guo, Yipin, et al.
Published: (2026)
SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference
by: Xia, Tian, et al.
Published: (2025)
by: Xia, Tian, et al.
Published: (2025)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
by: Fan, Ruibo, et al.
Published: (2026)
by: Fan, Ruibo, et al.
Published: (2026)
Revisiting Cache Freshness for Emerging Real-Time Applications
by: Mao, Ziming, et al.
Published: (2024)
by: Mao, Ziming, et al.
Published: (2024)
ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling
by: Yang, Yuchen, et al.
Published: (2026)
by: Yang, Yuchen, et al.
Published: (2026)
ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and Compression
by: Wang, Zirui, et al.
Published: (2025)
by: Wang, Zirui, et al.
Published: (2025)
gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters
by: Huang, Jiajun, et al.
Published: (2023)
by: Huang, Jiajun, et al.
Published: (2023)
Floating-Point Data Transformation for Lossless Compression
by: Jamalidinan, Samirasadat, et al.
Published: (2025)
by: Jamalidinan, Samirasadat, et al.
Published: (2025)
NCCLZ: Compression-Enabled GPU Collectives with Decoupled Quantization and Entropy Coding
by: Wang, Jiamin, et al.
Published: (2026)
by: Wang, Jiamin, et al.
Published: (2026)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
by: Griggs, Tyler, et al.
Published: (2024)
by: Griggs, Tyler, et al.
Published: (2024)
Towards Efficient and Practical GPU Multitasking in the Era of LLM
by: Xing, Jiarong, et al.
Published: (2025)
by: Xing, Jiarong, et al.
Published: (2025)
Unleashing Scalable Context Parallelism for Foundation Models Pre-Training via FCP
by: Zhao, Yilong, et al.
Published: (2026)
by: Zhao, Yilong, et al.
Published: (2026)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
by: Jiang, Xuanlin, et al.
Published: (2024)
by: Jiang, Xuanlin, et al.
Published: (2024)
Boosting Scientific Error-Bounded Lossy Compression through Optimized Synergistic Lossy-Lossless Orchestration
by: Wu, Shixun, et al.
Published: (2025)
by: Wu, Shixun, et al.
Published: (2025)
SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
by: Mao, Ziming, et al.
Published: (2024)
by: Mao, Ziming, et al.
Published: (2024)
DiffusionPipe: Training Large Diffusion Models with Efficient Pipelines
by: Tian, Ye, et al.
Published: (2024)
by: Tian, Ye, et al.
Published: (2024)
70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
On Optimizing the Communication of Model Parallelism
by: Zhuang, Yonghao, et al.
Published: (2022)
by: Zhuang, Yonghao, et al.
Published: (2022)
EXaCTz: Guaranteed Extremum Graph and Contour Tree Preservation for Distributed- and GPU-Parallel Lossy Compression
by: Li, Yuxiao, et al.
Published: (2026)
by: Li, Yuxiao, et al.
Published: (2026)
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
by: Liu, Xueshen, et al.
Published: (2026)
by: Liu, Xueshen, et al.
Published: (2026)
A Transverse-Read-assisted Fast Valid-Bits Collection in Stochastic Computing MACs for Energy-Efficient in-RTM DNNs
by: Wang, Jihe, et al.
Published: (2024)
by: Wang, Jihe, et al.
Published: (2024)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
by: Sojoodi, Amirhossein, et al.
Published: (2026)
by: Sojoodi, Amirhossein, et al.
Published: (2026)
Supercharging Federated Learning with Flower and NVIDIA FLARE
by: Roth, Holger R., et al.
Published: (2024)
by: Roth, Holger R., et al.
Published: (2024)
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
by: Duan, Jiangfei, et al.
Published: (2024)
by: Duan, Jiangfei, et al.
Published: (2024)
ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-Serving
by: Qiao, Yifan, et al.
Published: (2024)
by: Qiao, Yifan, et al.
Published: (2024)
AGILE: Lightweight and Efficient Asynchronous GPU-SSD Integration
by: Yang, Zhuoping, et al.
Published: (2025)
by: Yang, Zhuoping, et al.
Published: (2025)
Understanding GPU Triggering APIs for MPI+X Communication
by: Bridges, Patrick G., et al.
Published: (2024)
by: Bridges, Patrick G., et al.
Published: (2024)
Overcoming Memory Constraints in Quantum Circuit Simulation with a High-Fidelity Compression Framework
by: Zhang, Boyuan, et al.
Published: (2024)
by: Zhang, Boyuan, et al.
Published: (2024)
Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapping
by: Jiang, Chenyu, et al.
Published: (2024)
by: Jiang, Chenyu, et al.
Published: (2024)
GPU Volume Rendering with Hierarchical Compression Using VDB
by: Zellmann, Stefan, et al.
Published: (2025)
by: Zellmann, Stefan, et al.
Published: (2025)
RL over Commodity Networks: Overcoming the Bandwidth Barrier with Lossless Sparse Deltas
by: Ruan, Chaoyi, et al.
Published: (2026)
by: Ruan, Chaoyi, et al.
Published: (2026)
Locality-aware Fair Scheduling in LLM Serving
by: Cao, Shiyi, et al.
Published: (2025)
by: Cao, Shiyi, et al.
Published: (2025)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
by: Wang, Tianyu, et al.
Published: (2024)
by: Wang, Tianyu, et al.
Published: (2024)
Comprehensive Deadlock Prevention for GPU Collective Communication
by: Pan, Lichen, et al.
Published: (2023)
by: Pan, Lichen, et al.
Published: (2023)
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
by: Du, Kuntai, et al.
Published: (2025)
by: Du, Kuntai, et al.
Published: (2025)
Characterization-Guided GPU Fault Resilience in NVIDIA MPS
by: Liu, Rixin, et al.
Published: (2026)
by: Liu, Rixin, et al.
Published: (2026)
HAS-GPU: Efficient Hybrid Auto-scaling with Fine-grained GPU Allocation for SLO-aware Serverless Inferences
by: Gu, Jianfeng, et al.
Published: (2025)
by: Gu, Jianfeng, et al.
Published: (2025)
Similar Items
-
UCCL-EP: Portable Expert-Parallel Communication
by: Mao, Ziming, et al.
Published: (2025) -
ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training
by: Lin, Wenxiang, et al.
Published: (2026) -
Pie: Pooling CPU Memory for LLM Inference
by: Xu, Yi, et al.
Published: (2024) -
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
by: Guo, Yipin, et al.
Published: (2026) -
SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference
by: Xia, Tian, et al.
Published: (2025)