Poor Man's Training on MCUs: A Memory-Efficient Quantized Back-Propagation-Free Approach
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913575207960576 |
|---|---|
| author | Zhao, Yequan Li, Hai Young, Ian Zhang, Zheng |
| author_facet | Zhao, Yequan Li, Hai Young, Ian Zhang, Zheng |
| contents | Back propagation (BP) is the default solution for gradient computation in neural network training. However, implementing BP-based training on various edge devices such as FPGA, microcontrollers (MCUs), and analog computing platforms face multiple major challenges, such as the lack of hardware resources, long time-to-market, and dramatic errors in a low-precision setting. This paper presents a simple BP-free training scheme on an MCU, which makes edge training hardware design as easy as inference hardware design. We adopt a quantized zeroth-order method to estimate the gradients of quantized model parameters, which can overcome the error of a straight-through estimator in a low-precision BP scheme. We further employ a few dimension reduction methods (e.g., node perturbation, sparse training) to improve the convergence of zeroth-order training. Experiment results show that our BP-free training achieves comparable performance as BP-based training on adapting a pre-trained image classifier to various corrupted data on resource-constrained edge devices (e.g., an MCU with 1024-KB SRAM for dense full-model training, or an MCU with 256-KB SRAM for sparse training). This method is most suitable for application scenarios where memory cost and time-to-market are the major concerns, but longer latency can be tolerated. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_05873 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Poor Man's Training on MCUs: A Memory-Efficient Quantized Back-Propagation-Free Approach Zhao, Yequan Li, Hai Young, Ian Zhang, Zheng Machine Learning Computer Vision and Pattern Recognition Distributed, Parallel, and Cluster Computing Neural and Evolutionary Computing I.2; C.3 Back propagation (BP) is the default solution for gradient computation in neural network training. However, implementing BP-based training on various edge devices such as FPGA, microcontrollers (MCUs), and analog computing platforms face multiple major challenges, such as the lack of hardware resources, long time-to-market, and dramatic errors in a low-precision setting. This paper presents a simple BP-free training scheme on an MCU, which makes edge training hardware design as easy as inference hardware design. We adopt a quantized zeroth-order method to estimate the gradients of quantized model parameters, which can overcome the error of a straight-through estimator in a low-precision BP scheme. We further employ a few dimension reduction methods (e.g., node perturbation, sparse training) to improve the convergence of zeroth-order training. Experiment results show that our BP-free training achieves comparable performance as BP-based training on adapting a pre-trained image classifier to various corrupted data on resource-constrained edge devices (e.g., an MCU with 1024-KB SRAM for dense full-model training, or an MCU with 256-KB SRAM for sparse training). This method is most suitable for application scenarios where memory cost and time-to-market are the major concerns, but longer latency can be tolerated. |
| title | Poor Man's Training on MCUs: A Memory-Efficient Quantized Back-Propagation-Free Approach |
| topic | Machine Learning Computer Vision and Pattern Recognition Distributed, Parallel, and Cluster Computing Neural and Evolutionary Computing I.2; C.3 |
| url | https://arxiv.org/abs/2411.05873 |