Accelerating TinyML Inference on Microcontrollers through Approximate Kernels

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Armeniakos, Giorgos, Mentzos, Georgios, Soudris, Dimitrios
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929515064721408
author Armeniakos, Giorgos
Mentzos, Georgios
Soudris, Dimitrios
author_facet Armeniakos, Giorgos
Mentzos, Georgios
Soudris, Dimitrios
contents The rapid growth of microcontroller-based IoT devices has opened up numerous applications, from smart manufacturing to personalized healthcare. Despite the widespread adoption of energy-efficient microcontroller units (MCUs) in the Tiny Machine Learning (TinyML) domain, they still face significant limitations in terms of performance and memory (RAM, Flash). In this work, we combine approximate computing and software kernel design to accelerate the inference of approximate CNN models on MCUs. Our kernel-based approximation framework firstly unpacks the operands of each convolution layer and then conducts an offline calculation to determine the significance of each operand. Subsequently, through a design space exploration, it employs a computation skipping approximation strategy based on the calculated significance. Our evaluation on an STM32-Nucleo board and 2 popular CNNs trained on the CIFAR-10 dataset shows that, compared to state-of-the-art exact inference, our Pareto optimal solutions can feature on average 21% latency reduction with no degradation in Top-1 classification accuracy, while for lower accuracy requirements, the corresponding reduction becomes even more pronounced.
format Preprint
id arxiv_https___arxiv_org_abs_2409_16815
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Accelerating TinyML Inference on Microcontrollers through Approximate Kernels
Armeniakos, Giorgos
Mentzos, Georgios
Soudris, Dimitrios
Machine Learning
The rapid growth of microcontroller-based IoT devices has opened up numerous applications, from smart manufacturing to personalized healthcare. Despite the widespread adoption of energy-efficient microcontroller units (MCUs) in the Tiny Machine Learning (TinyML) domain, they still face significant limitations in terms of performance and memory (RAM, Flash). In this work, we combine approximate computing and software kernel design to accelerate the inference of approximate CNN models on MCUs. Our kernel-based approximation framework firstly unpacks the operands of each convolution layer and then conducts an offline calculation to determine the significance of each operand. Subsequently, through a design space exploration, it employs a computation skipping approximation strategy based on the calculated significance. Our evaluation on an STM32-Nucleo board and 2 popular CNNs trained on the CIFAR-10 dataset shows that, compared to state-of-the-art exact inference, our Pareto optimal solutions can feature on average 21% latency reduction with no degradation in Top-1 classification accuracy, while for lower accuracy requirements, the corresponding reduction becomes even more pronounced.
title Accelerating TinyML Inference on Microcontrollers through Approximate Kernels
topic Machine Learning
url https://arxiv.org/abs/2409.16815