Benchmarking Ultra-Low-Power $μ$NPUs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Millar, Josh, Huang, Yushan, Sethi, Sarab, Haddadi, Hamed, Madhavapeddy, Anil
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915587818520576
author Millar, Josh
Huang, Yushan
Sethi, Sarab
Haddadi, Hamed
Madhavapeddy, Anil
author_facet Millar, Josh
Huang, Yushan
Sethi, Sarab
Haddadi, Hamed
Madhavapeddy, Anil
contents Efficient on-device neural network (NN) inference offers predictable latency, improved privacy and reliability, and lower operating costs for vendors than cloud-based inference. This has sparked recent development of microcontroller-scale NN accelerators, also known as neural processing units ($μ$NPUs), designed specifically for ultra-low-power applications. We present the first comparative evaluation of a number of commercially-available $μ$NPUs, including the first independent benchmarks for multiple platforms. To ensure fairness, we develop and open-source a model compilation pipeline supporting consistent benchmarking of quantized models across diverse microcontroller hardware. Our resulting analysis uncovers both expected performance trends as well as surprising disparities between hardware specifications and actual performance, including certain $μ$NPUs exhibiting unexpected scaling behaviors with model complexity. This work provides a foundation for ongoing evaluation of $μ$NPU platforms, alongside offering practical insights for both hardware and software developers in this rapidly evolving space.
format Preprint
id arxiv_https___arxiv_org_abs_2503_22567
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking Ultra-Low-Power $μ$NPUs
Millar, Josh
Huang, Yushan
Sethi, Sarab
Haddadi, Hamed
Madhavapeddy, Anil
Machine Learning
Hardware Architecture
Efficient on-device neural network (NN) inference offers predictable latency, improved privacy and reliability, and lower operating costs for vendors than cloud-based inference. This has sparked recent development of microcontroller-scale NN accelerators, also known as neural processing units ($μ$NPUs), designed specifically for ultra-low-power applications. We present the first comparative evaluation of a number of commercially-available $μ$NPUs, including the first independent benchmarks for multiple platforms. To ensure fairness, we develop and open-source a model compilation pipeline supporting consistent benchmarking of quantized models across diverse microcontroller hardware. Our resulting analysis uncovers both expected performance trends as well as surprising disparities between hardware specifications and actual performance, including certain $μ$NPUs exhibiting unexpected scaling behaviors with model complexity. This work provides a foundation for ongoing evaluation of $μ$NPU platforms, alongside offering practical insights for both hardware and software developers in this rapidly evolving space.
title Benchmarking Ultra-Low-Power $μ$NPUs
topic Machine Learning
Hardware Architecture
url https://arxiv.org/abs/2503.22567