Saved in:
Bibliographic Details
Main Authors: Li, Shiming, Mottola, Luca, Yao, Yuan, Kaxiras, Stefanos
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.05347
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910033990647808
author Li, Shiming
Mottola, Luca
Yao, Yuan
Kaxiras, Stefanos
author_facet Li, Shiming
Mottola, Luca
Yao, Yuan
Kaxiras, Stefanos
contents Quantized CNN inference on ultra-low-power MCUs incurs unnecessary computations in neurons that produce saturated output values. These values are too extreme and are eventually clamped to the boundaries allowed by the neuron. Often times, the neuron can save time by only producing a value that is extreme enough to lead to the clamped result, instead of completing the computation, yet without introducing any error. Based on this, we present saturation-aware convolution: an inference technique whereby we alter the order of computations in convolution kernels to induce earlier saturation, and value checks are inserted to omit unnecessary computations when the intermediate result is sufficiently extreme. Our experimental results display up to 24% inference time saving on a Cortex-M0+ MCU, with zero impact on accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05347
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient CNN Inference on Ultra-Low-Power MCUs via Saturation-Aware Convolution
Li, Shiming
Mottola, Luca
Yao, Yuan
Kaxiras, Stefanos
Systems and Control
Quantized CNN inference on ultra-low-power MCUs incurs unnecessary computations in neurons that produce saturated output values. These values are too extreme and are eventually clamped to the boundaries allowed by the neuron. Often times, the neuron can save time by only producing a value that is extreme enough to lead to the clamped result, instead of completing the computation, yet without introducing any error. Based on this, we present saturation-aware convolution: an inference technique whereby we alter the order of computations in convolution kernels to induce earlier saturation, and value checks are inserted to omit unnecessary computations when the intermediate result is sufficiently extreme. Our experimental results display up to 24% inference time saving on a Cortex-M0+ MCU, with zero impact on accuracy.
title Efficient CNN Inference on Ultra-Low-Power MCUs via Saturation-Aware Convolution
topic Systems and Control
url https://arxiv.org/abs/2511.05347