Perturbation-efficient Zeroth-order Optimization for Hardware-friendly On-device Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Qitao, Chang, Sung-En, Xia, Rui, Ji, Huidong, Yang, Chence, Zhang, Ci, Liu, Jun, Zhan, Zheng, Fang, Zhenman, Zou, Zhou, Wang, Yanzhi, Lu, Jin, Yuan, Geng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911075507634176
author Tan, Qitao
Chang, Sung-En
Xia, Rui
Ji, Huidong
Yang, Chence
Zhang, Ci
Liu, Jun
Zhan, Zheng
Fang, Zhenman
Zou, Zhou
Wang, Yanzhi
Lu, Jin
Yuan, Geng
author_facet Tan, Qitao
Chang, Sung-En
Xia, Rui
Ji, Huidong
Yang, Chence
Zhang, Ci
Liu, Jun
Zhan, Zheng
Fang, Zhenman
Zou, Zhou
Wang, Yanzhi
Lu, Jin
Yuan, Geng
contents Zeroth-order (ZO) optimization is an emerging deep neural network (DNN) training paradigm that offers computational simplicity and memory savings. However, this seemingly promising approach faces a significant and long-ignored challenge. ZO requires generating a substantial number of Gaussian random numbers, which poses significant difficulties and even makes it infeasible for hardware platforms, such as FPGAs and ASICs. In this paper, we identify this critical issue, which arises from the mismatch between algorithm and hardware designers. To address this issue, we proposed PeZO, a perturbation-efficient ZO framework. Specifically, we design random number reuse strategies to significantly reduce the demand for random number generation and introduce a hardware-friendly adaptive scaling method to replace the costly Gaussian distribution with a uniform distribution. Our experiments show that PeZO reduces the required LUTs and FFs for random number generation by 48.6\% and 12.7\%, and saves at maximum 86\% power consumption, all without compromising training performance, making ZO optimization feasible for on-device training. To the best of our knowledge, we are the first to explore the potential of on-device ZO optimization, providing valuable insights for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2504_20314
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Perturbation-efficient Zeroth-order Optimization for Hardware-friendly On-device Training
Tan, Qitao
Chang, Sung-En
Xia, Rui
Ji, Huidong
Yang, Chence
Zhang, Ci
Liu, Jun
Zhan, Zheng
Fang, Zhenman
Zou, Zhou
Wang, Yanzhi
Lu, Jin
Yuan, Geng
Machine Learning
Artificial Intelligence
Zeroth-order (ZO) optimization is an emerging deep neural network (DNN) training paradigm that offers computational simplicity and memory savings. However, this seemingly promising approach faces a significant and long-ignored challenge. ZO requires generating a substantial number of Gaussian random numbers, which poses significant difficulties and even makes it infeasible for hardware platforms, such as FPGAs and ASICs. In this paper, we identify this critical issue, which arises from the mismatch between algorithm and hardware designers. To address this issue, we proposed PeZO, a perturbation-efficient ZO framework. Specifically, we design random number reuse strategies to significantly reduce the demand for random number generation and introduce a hardware-friendly adaptive scaling method to replace the costly Gaussian distribution with a uniform distribution. Our experiments show that PeZO reduces the required LUTs and FFs for random number generation by 48.6\% and 12.7\%, and saves at maximum 86\% power consumption, all without compromising training performance, making ZO optimization feasible for on-device training. To the best of our knowledge, we are the first to explore the potential of on-device ZO optimization, providing valuable insights for future research.
title Perturbation-efficient Zeroth-order Optimization for Hardware-friendly On-device Training
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2504.20314