eIQ Neutron: Redefining Edge-AI Inference with Integrated NPU and Compiler Innovations
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916955545403392 |
|---|---|
| author | Bamberg, Lennart Minnella, Filippo Bosio, Roberto Ottati, Fabrizio Wang, Yuebin Lee, Jongmin Lavagno, Luciano Fuks, Adam |
| author_facet | Bamberg, Lennart Minnella, Filippo Bosio, Roberto Ottati, Fabrizio Wang, Yuebin Lee, Jongmin Lavagno, Luciano Fuks, Adam |
| contents | Neural Processing Units (NPUs) are key to enabling efficient AI inference in resource-constrained edge environments. While peak tera operations per second (TOPS) is often used to gauge performance, it poorly reflects real-world performance and typically rather correlates with higher silicon cost. To address this, architects must focus on maximizing compute utilization, without sacrificing flexibility. This paper presents the eIQ Neutron efficient-NPU, integrated into a commercial flagship MPU, alongside co-designed compiler algorithms. The architecture employs a flexible, data-driven design, while the compiler uses a constrained programming approach to optimize compute and data movement based on workload characteristics. Compared to the leading embedded NPU and compiler stack, our solution achieves an average speedup of 1.8x (4x peak) at equal TOPS and memory resources across standard AI-benchmarks. Even against NPUs with double the compute and memory resources, Neutron delivers up to 3.3x higher performance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_14388 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | eIQ Neutron: Redefining Edge-AI Inference with Integrated NPU and Compiler Innovations Bamberg, Lennart Minnella, Filippo Bosio, Roberto Ottati, Fabrizio Wang, Yuebin Lee, Jongmin Lavagno, Luciano Fuks, Adam Hardware Architecture Artificial Intelligence Machine Learning Neural Processing Units (NPUs) are key to enabling efficient AI inference in resource-constrained edge environments. While peak tera operations per second (TOPS) is often used to gauge performance, it poorly reflects real-world performance and typically rather correlates with higher silicon cost. To address this, architects must focus on maximizing compute utilization, without sacrificing flexibility. This paper presents the eIQ Neutron efficient-NPU, integrated into a commercial flagship MPU, alongside co-designed compiler algorithms. The architecture employs a flexible, data-driven design, while the compiler uses a constrained programming approach to optimize compute and data movement based on workload characteristics. Compared to the leading embedded NPU and compiler stack, our solution achieves an average speedup of 1.8x (4x peak) at equal TOPS and memory resources across standard AI-benchmarks. Even against NPUs with double the compute and memory resources, Neutron delivers up to 3.3x higher performance. |
| title | eIQ Neutron: Redefining Edge-AI Inference with Integrated NPU and Compiler Innovations |
| topic | Hardware Architecture Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2509.14388 |