EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yi, Qingao, Duan, Jiaang, Hu, Hanwen, Hua, Qin, Zhao, Haiyan, Qian, Shiyou, Yang, Dingyu, Cao, Jian, Tang, Jinghua, Yu, Yinghao, Liao, Chenzhi, Wang, Kangjin, Zhang, Liping
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915677707698176
author Yi, Qingao
Duan, Jiaang
Hu, Hanwen
Hua, Qin
Zhao, Haiyan
Qian, Shiyou
Yang, Dingyu
Cao, Jian
Tang, Jinghua
Yu, Yinghao
Liao, Chenzhi
Wang, Kangjin
Zhang, Liping
author_facet Yi, Qingao
Duan, Jiaang
Hu, Hanwen
Hua, Qin
Zhao, Haiyan
Qian, Shiyou
Yang, Dingyu
Cao, Jian
Tang, Jinghua
Yu, Yinghao
Liao, Chenzhi
Wang, Kangjin
Zhang, Liping
contents Training large language models (LLMs) poses significant challenges regarding computational resources and memory capacity. Although distributed training techniques help mitigate these issues, they still suffer from considerable communication overhead. Existing approaches primarily rely on static gradient compression to enhance communication efficiency; however, these methods neglect the dynamic nature of evolving gradients during training, leading to performance degradation. Accelerating LLM training via compression without sacrificing performance remains a challenge. In this paper, we propose an entropy-driven dynamic gradient compression framework called EDGC. The core concept is to adjust the compression rate during LLM training based on the evolving trends of gradient entropy, taking into account both compression efficiency and error. EDGC consists of three key components.First, it employs a down-sampling method to efficiently estimate gradient entropy, reducing computation overhead. Second, it establishes a theoretical model linking compression rate with gradient entropy, enabling more informed compression decisions. Lastly, a window-based adjustment mechanism dynamically adapts the compression rate across pipeline stages, improving communication efficiency and maintaining model performance. We implemented EDGC on a 32-NVIDIA-V100 cluster and a 64-NVIDIA-H100 cluster to train GPT2-2.5B and GPT2-12.1B, respectively. The results show that EDGC significantly reduces communication latency and training time by up to 46.45% and 16.13% while preserving LLM accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10333
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
Yi, Qingao
Duan, Jiaang
Hu, Hanwen
Hua, Qin
Zhao, Haiyan
Qian, Shiyou
Yang, Dingyu
Cao, Jian
Tang, Jinghua
Yu, Yinghao
Liao, Chenzhi
Wang, Kangjin
Zhang, Liping
Machine Learning
Artificial Intelligence
Performance
Training large language models (LLMs) poses significant challenges regarding computational resources and memory capacity. Although distributed training techniques help mitigate these issues, they still suffer from considerable communication overhead. Existing approaches primarily rely on static gradient compression to enhance communication efficiency; however, these methods neglect the dynamic nature of evolving gradients during training, leading to performance degradation. Accelerating LLM training via compression without sacrificing performance remains a challenge. In this paper, we propose an entropy-driven dynamic gradient compression framework called EDGC. The core concept is to adjust the compression rate during LLM training based on the evolving trends of gradient entropy, taking into account both compression efficiency and error. EDGC consists of three key components.First, it employs a down-sampling method to efficiently estimate gradient entropy, reducing computation overhead. Second, it establishes a theoretical model linking compression rate with gradient entropy, enabling more informed compression decisions. Lastly, a window-based adjustment mechanism dynamically adapts the compression rate across pipeline stages, improving communication efficiency and maintaining model performance. We implemented EDGC on a 32-NVIDIA-V100 cluster and a 64-NVIDIA-H100 cluster to train GPT2-2.5B and GPT2-12.1B, respectively. The results show that EDGC significantly reduces communication latency and training time by up to 46.45% and 16.13% while preserving LLM accuracy.
title EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
topic Machine Learning
Artificial Intelligence
Performance
url https://arxiv.org/abs/2511.10333