On-Device Training of Fully Quantized Deep Neural Networks on Cortex-M Microcontrollers

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Deutel, Mark, Hannig, Frank, Mutschler, Christopher, Teich, Jürgen
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912082817974272
author Deutel, Mark
Hannig, Frank
Mutschler, Christopher
Teich, Jürgen
author_facet Deutel, Mark
Hannig, Frank
Mutschler, Christopher
Teich, Jürgen
contents On-device training of DNNs allows models to adapt and fine-tune to newly collected data or changing domains while deployed on microcontroller units (MCUs). However, DNN training is a resource-intensive task, making the implementation and execution of DNN training algorithms on MCUs challenging due to low processor speeds, constrained throughput, limited floating-point support, and memory constraints. In this work, we explore on-device training of DNNs for Cortex-M MCUs. We present a method that enables efficient training of DNNs completely in place on the MCU using fully quantized training (FQT) and dynamic partial gradient updates. We demonstrate the feasibility of our approach on multiple vision and time-series datasets and provide insights into the tradeoff between training accuracy, memory overhead, energy, and latency on real hardware.
format Preprint
id arxiv_https___arxiv_org_abs_2407_10734
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On-Device Training of Fully Quantized Deep Neural Networks on Cortex-M Microcontrollers
Deutel, Mark
Hannig, Frank
Mutschler, Christopher
Teich, Jürgen
Machine Learning
Artificial Intelligence
On-device training of DNNs allows models to adapt and fine-tune to newly collected data or changing domains while deployed on microcontroller units (MCUs). However, DNN training is a resource-intensive task, making the implementation and execution of DNN training algorithms on MCUs challenging due to low processor speeds, constrained throughput, limited floating-point support, and memory constraints. In this work, we explore on-device training of DNNs for Cortex-M MCUs. We present a method that enables efficient training of DNNs completely in place on the MCU using fully quantized training (FQT) and dynamic partial gradient updates. We demonstrate the feasibility of our approach on multiple vision and time-series datasets and provide insights into the tradeoff between training accuracy, memory overhead, energy, and latency on real hardware.
title On-Device Training of Fully Quantized Deep Neural Networks on Cortex-M Microcontrollers
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2407.10734