AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Wenxiang, Huang, Juntao, Zhang, Luhan, Li, Laili, Bao, Xiang, Zhang, Mengyang, Wang, Bing, Shi, Shaohuai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules
von: Pan, Xinglin, et al.
Veröffentlicht: (2024)
von: Pan, Xinglin, et al.
Veröffentlicht: (2024)
ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training
von: Lin, Wenxiang, et al.
Veröffentlicht: (2026)
von: Lin, Wenxiang, et al.
Veröffentlicht: (2026)
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
von: Lin, Wenxiang, et al.
Veröffentlicht: (2025)
von: Lin, Wenxiang, et al.
Veröffentlicht: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
QSync: Quantization-Minimized Synchronous Distributed Training Across Hybrid Devices
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
DreamDDP: Accelerating Data Parallel Distributed LLM Training with Layer-wise Scheduled Partial Synchronization
von: Tang, Zhenheng, et al.
Veröffentlicht: (2025)
von: Tang, Zhenheng, et al.
Veröffentlicht: (2025)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
von: Guo, Cong, et al.
Veröffentlicht: (2024)
von: Guo, Cong, et al.
Veröffentlicht: (2024)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
On the Performance and Memory Footprint of Distributed Training: An Empirical Study on Transformers
von: Lu, Zhengxian, et al.
Veröffentlicht: (2024)
von: Lu, Zhengxian, et al.
Veröffentlicht: (2024)
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
von: Zhang, Han, et al.
Veröffentlicht: (2026)
von: Zhang, Han, et al.
Veröffentlicht: (2026)
Optimizing Federated Learning in the Era of LLMs: Message Quantization and Streaming
von: Xu, Ziyue, et al.
Veröffentlicht: (2025)
von: Xu, Ziyue, et al.
Veröffentlicht: (2025)
Efficient Distributed MLLM Training with Cornstarch
von: Jang, Insu, et al.
Veröffentlicht: (2025)
von: Jang, Insu, et al.
Veröffentlicht: (2025)
Scheduling Deep Learning Jobs in Multi-Tenant GPU Clusters via Wise Resource Sharing
von: Luo, Yizhou, et al.
Veröffentlicht: (2024)
von: Luo, Yizhou, et al.
Veröffentlicht: (2024)
CondenseGraph: Communication-Efficient Distributed GNN Training via On-the-Fly Graph Condensation
von: Zhang, Zizhao, et al.
Veröffentlicht: (2026)
von: Zhang, Zizhao, et al.
Veröffentlicht: (2026)
Efficient and Portable Support for Overdecomposition on Distributed Memory GPGPU Platforms
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
Communication-Efficient Distributed Learning via Sparse and Adaptive Stochastic Gradient
von: Deng, Xiaoge, et al.
Veröffentlicht: (2021)
von: Deng, Xiaoge, et al.
Veröffentlicht: (2021)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
von: Wang, Yuxin, et al.
Veröffentlicht: (2023)
MTGenRec: An Efficient Distributed Training System for Generative Recommendation Models in Meituan
von: Wang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Wang, Yuxiang, et al.
Veröffentlicht: (2025)
DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training
von: Wang, Zhixin, et al.
Veröffentlicht: (2025)
von: Wang, Zhixin, et al.
Veröffentlicht: (2025)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yaozheng, et al.
Veröffentlicht: (2025)
FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
von: Gao, Yunqi, et al.
Veröffentlicht: (2025)
von: Gao, Yunqi, et al.
Veröffentlicht: (2025)
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation
von: Chen, Fahao, et al.
Veröffentlicht: (2024)
von: Chen, Fahao, et al.
Veröffentlicht: (2024)
Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
von: He, Guoliang, et al.
Veröffentlicht: (2025)
von: He, Guoliang, et al.
Veröffentlicht: (2025)
DawnPiper: A Memory-scablable Pipeline Parallel Training Framework
von: Peng, Xuan, et al.
Veröffentlicht: (2025)
von: Peng, Xuan, et al.
Veröffentlicht: (2025)
Clock Distribution with Gradient TRIX
von: Lenzen, Christoph, et al.
Veröffentlicht: (2023)
von: Lenzen, Christoph, et al.
Veröffentlicht: (2023)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMs
von: Lee, Haeun, et al.
Veröffentlicht: (2025)
von: Lee, Haeun, et al.
Veröffentlicht: (2025)
Memory Efficient and Staleness Free Pipeline Parallel DNN Training Framework with Improved Convergence Speed
von: Dutta, Ankita, et al.
Veröffentlicht: (2025)
von: Dutta, Ankita, et al.
Veröffentlicht: (2025)
AME: An Efficient Heterogeneous Agentic Memory Engine for Smartphones
von: Zhao, Xinkui, et al.
Veröffentlicht: (2025)
von: Zhao, Xinkui, et al.
Veröffentlicht: (2025)
HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments
von: He, Yongjun, et al.
Veröffentlicht: (2025)
von: He, Yongjun, et al.
Veröffentlicht: (2025)
TiMePReSt: Time and Memory Efficient Pipeline Parallel DNN Training with Removed Staleness
von: Dutta, Ankita, et al.
Veröffentlicht: (2024)
von: Dutta, Ankita, et al.
Veröffentlicht: (2024)
Byzantine-Robust and Communication-Efficient Distributed Training: Compressive and Cyclic Gradient Coding
von: Li, Chengxi, et al.
Veröffentlicht: (2026)
von: Li, Chengxi, et al.
Veröffentlicht: (2026)
Distributed Order Recording Techniques for Efficient Record-and-Replay of Multi-threaded Programs
von: Fu, Xiang, et al.
Veröffentlicht: (2026)
von: Fu, Xiang, et al.
Veröffentlicht: (2026)
RapidGNN: Communication Efficient Large-Scale Distributed Training of Graph Neural Networks
von: Niam, Arefin, et al.
Veröffentlicht: (2025)
von: Niam, Arefin, et al.
Veröffentlicht: (2025)
GreenDyGNN: Runtime-Adaptive Energy-Efficient Communication for Distributed GNN Training
von: Niam, Arefin, et al.
Veröffentlicht: (2026)
von: Niam, Arefin, et al.
Veröffentlicht: (2026)
Biased Compression in Gradient Coding for Distributed Learning
von: Li, Chengxi, et al.
Veröffentlicht: (2026)
von: Li, Chengxi, et al.
Veröffentlicht: (2026)
DiFache: Efficient and Scalable Caching on Disaggregated Memory using Decentralized Coherence
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules
von: Pan, Xinglin, et al.
Veröffentlicht: (2024) -
ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training
von: Lin, Wenxiang, et al.
Veröffentlicht: (2026) -
HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
von: Lin, Wenxiang, et al.
Veröffentlicht: (2025) -
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025) -
QSync: Quantization-Minimized Synchronous Distributed Training Across Hybrid Devices
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)