eIQ Neutron: Redefining Edge-AI Inference with Integrated NPU and Compiler Innovations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bamberg, Lennart, Minnella, Filippo, Bosio, Roberto, Ottati, Fabrizio, Wang, Yuebin, Lee, Jongmin, Lavagno, Luciano, Fuks, Adam
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916955545403392
author Bamberg, Lennart
Minnella, Filippo
Bosio, Roberto
Ottati, Fabrizio
Wang, Yuebin
Lee, Jongmin
Lavagno, Luciano
Fuks, Adam
author_facet Bamberg, Lennart
Minnella, Filippo
Bosio, Roberto
Ottati, Fabrizio
Wang, Yuebin
Lee, Jongmin
Lavagno, Luciano
Fuks, Adam
contents Neural Processing Units (NPUs) are key to enabling efficient AI inference in resource-constrained edge environments. While peak tera operations per second (TOPS) is often used to gauge performance, it poorly reflects real-world performance and typically rather correlates with higher silicon cost. To address this, architects must focus on maximizing compute utilization, without sacrificing flexibility. This paper presents the eIQ Neutron efficient-NPU, integrated into a commercial flagship MPU, alongside co-designed compiler algorithms. The architecture employs a flexible, data-driven design, while the compiler uses a constrained programming approach to optimize compute and data movement based on workload characteristics. Compared to the leading embedded NPU and compiler stack, our solution achieves an average speedup of 1.8x (4x peak) at equal TOPS and memory resources across standard AI-benchmarks. Even against NPUs with double the compute and memory resources, Neutron delivers up to 3.3x higher performance.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14388
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle eIQ Neutron: Redefining Edge-AI Inference with Integrated NPU and Compiler Innovations
Bamberg, Lennart
Minnella, Filippo
Bosio, Roberto
Ottati, Fabrizio
Wang, Yuebin
Lee, Jongmin
Lavagno, Luciano
Fuks, Adam
Hardware Architecture
Artificial Intelligence
Machine Learning
Neural Processing Units (NPUs) are key to enabling efficient AI inference in resource-constrained edge environments. While peak tera operations per second (TOPS) is often used to gauge performance, it poorly reflects real-world performance and typically rather correlates with higher silicon cost. To address this, architects must focus on maximizing compute utilization, without sacrificing flexibility. This paper presents the eIQ Neutron efficient-NPU, integrated into a commercial flagship MPU, alongside co-designed compiler algorithms. The architecture employs a flexible, data-driven design, while the compiler uses a constrained programming approach to optimize compute and data movement based on workload characteristics. Compared to the leading embedded NPU and compiler stack, our solution achieves an average speedup of 1.8x (4x peak) at equal TOPS and memory resources across standard AI-benchmarks. Even against NPUs with double the compute and memory resources, Neutron delivers up to 3.3x higher performance.
title eIQ Neutron: Redefining Edge-AI Inference with Integrated NPU and Compiler Innovations
topic Hardware Architecture
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.14388