Rethinking the Potential of Layer Freezing for Efficient DNN Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Chence, Zhang, Ci, Lu, Lei, Tan, Qitao, Li, Sheng, Li, Ao, Tang, Xulong, Huang, Shaoyi, Wang, Jinzhen, Li, Guoming, Li, Jundong, Zhai, Xiaoming, Lu, Jin, Yuan, Geng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913999224832000
author Yang, Chence
Zhang, Ci
Lu, Lei
Tan, Qitao
Li, Sheng
Li, Ao
Tang, Xulong
Huang, Shaoyi
Wang, Jinzhen
Li, Guoming
Li, Jundong
Zhai, Xiaoming
Lu, Jin
Yuan, Geng
author_facet Yang, Chence
Zhang, Ci
Lu, Lei
Tan, Qitao
Li, Sheng
Li, Ao
Tang, Xulong
Huang, Shaoyi
Wang, Jinzhen
Li, Guoming
Li, Jundong
Zhai, Xiaoming
Lu, Jin
Yuan, Geng
contents With the growing size of deep neural networks and datasets, the computational costs of training have significantly increased. The layer-freezing technique has recently attracted great attention as a promising method to effectively reduce the cost of network training. However, in traditional layer-freezing methods, frozen layers are still required for forward propagation to generate feature maps for unfrozen layers, limiting the reduction of computation costs. To overcome this, prior works proposed a hypothetical solution, which caches feature maps from frozen layers as a new dataset, allowing later layers to train directly on stored feature maps. While this approach appears to be straightforward, it presents several major challenges that are severely overlooked by prior literature, such as how to effectively apply augmentations to feature maps and the substantial storage overhead introduced. If these overlooked challenges are not addressed, the performance of the caching method will be severely impacted and even make it infeasible. This paper is the first to comprehensively explore these challenges and provides a systematic solution. To improve training accuracy, we propose \textit{similarity-aware channel augmentation}, which caches channels with high augmentation sensitivity with a minimum additional storage cost. To mitigate storage overhead, we incorporate lossy data compression into layer freezing and design a \textit{progressive compression} strategy, which increases compression rates as more layers are frozen, effectively reducing storage costs. Finally, our solution achieves significant reductions in training cost while maintaining model accuracy, with a minor time overhead. Additionally, we conduct a comprehensive evaluation of freezing and compression strategies, providing insights into optimizing their application for efficient DNN training.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15033
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rethinking the Potential of Layer Freezing for Efficient DNN Training
Yang, Chence
Zhang, Ci
Lu, Lei
Tan, Qitao
Li, Sheng
Li, Ao
Tang, Xulong
Huang, Shaoyi
Wang, Jinzhen
Li, Guoming
Li, Jundong
Zhai, Xiaoming
Lu, Jin
Yuan, Geng
Machine Learning
With the growing size of deep neural networks and datasets, the computational costs of training have significantly increased. The layer-freezing technique has recently attracted great attention as a promising method to effectively reduce the cost of network training. However, in traditional layer-freezing methods, frozen layers are still required for forward propagation to generate feature maps for unfrozen layers, limiting the reduction of computation costs. To overcome this, prior works proposed a hypothetical solution, which caches feature maps from frozen layers as a new dataset, allowing later layers to train directly on stored feature maps. While this approach appears to be straightforward, it presents several major challenges that are severely overlooked by prior literature, such as how to effectively apply augmentations to feature maps and the substantial storage overhead introduced. If these overlooked challenges are not addressed, the performance of the caching method will be severely impacted and even make it infeasible. This paper is the first to comprehensively explore these challenges and provides a systematic solution. To improve training accuracy, we propose \textit{similarity-aware channel augmentation}, which caches channels with high augmentation sensitivity with a minimum additional storage cost. To mitigate storage overhead, we incorporate lossy data compression into layer freezing and design a \textit{progressive compression} strategy, which increases compression rates as more layers are frozen, effectively reducing storage costs. Finally, our solution achieves significant reductions in training cost while maintaining model accuracy, with a minor time overhead. Additionally, we conduct a comprehensive evaluation of freezing and compression strategies, providing insights into optimizing their application for efficient DNN training.
title Rethinking the Potential of Layer Freezing for Efficient DNN Training
topic Machine Learning
url https://arxiv.org/abs/2508.15033