eIQ Neutron: Redefining Edge-AI Inference with Integrated NPU and Compiler Innovations
Fuente:
arXiv
Saved in:
| Main Authors: | Bamberg, Lennart, Minnella, Filippo, Bosio, Roberto, Ottati, Fabrizio, Wang, Yuebin, Lee, Jongmin, Lavagno, Luciano, Fuks, Adam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SILVIA: Automated Superword-Level Parallelism Exploitation via HLS-Specific LLVM Passes for Compute-Intensive FPGA Accelerators
by: Brignone, Giovanni, et al.
Published: (2024)
by: Brignone, Giovanni, et al.
Published: (2024)
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
by: Mueller, Lion, et al.
Published: (2025)
by: Mueller, Lion, et al.
Published: (2025)
A DSP shared is a DSP earned: HLS Task-Level Multi-Pumping for High-Performance Low-Resource Designs
by: Brignone, Giovanni, et al.
Published: (2023)
by: Brignone, Giovanni, et al.
Published: (2023)
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
by: Chen, Yuzong, et al.
Published: (2025)
by: Chen, Yuzong, et al.
Published: (2025)
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
by: Heo, Guseul, et al.
Published: (2024)
by: Heo, Guseul, et al.
Published: (2024)
IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
by: Seo, Minseok, et al.
Published: (2024)
by: Seo, Minseok, et al.
Published: (2024)
Just TestIt! An SBST Approach To Automate System-Integration Testing
by: Terzano, Tommaso, et al.
Published: (2025)
by: Terzano, Tommaso, et al.
Published: (2025)
To Spike or Not To Spike: A Digital Hardware Perspective on Deep Learning Acceleration
by: Ottati, Fabrizio, et al.
Published: (2023)
by: Ottati, Fabrizio, et al.
Published: (2023)
EONSim: An NPU Simulator for On-Chip Memory and Embedding Vector Operations
by: Choi, Sangun, et al.
Published: (2025)
by: Choi, Sangun, et al.
Published: (2025)
PowerFlow-DNN: Compiler-Directed Fine-Grained Power Orchestration for End-to-End Edge AI Inference
by: Chen, Paul, et al.
Published: (2026)
by: Chen, Paul, et al.
Published: (2026)
VUSA: Virtually Upscaled Systolic Array Architecture to Exploit Unstructured Sparsity in AI Acceleration
by: Helal, Shereef, et al.
Published: (2025)
by: Helal, Shereef, et al.
Published: (2025)
From Characterization to Microarchitecture: Designing an Elegant and Reliable BFP-Based NPU
by: Zhang, Jie, et al.
Published: (2026)
by: Zhang, Jie, et al.
Published: (2026)
Strix: Re-thinking NPU Reliability from a System Perspective
by: Guan, Jiapeng, et al.
Published: (2026)
by: Guan, Jiapeng, et al.
Published: (2026)
Implementation and Optimization of HQC Decoding on NPU-Integrated Devices
by: Chau, Vu Minh, et al.
Published: (2026)
by: Chau, Vu Minh, et al.
Published: (2026)
ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
by: Ham, Hyungkyu, et al.
Published: (2024)
by: Ham, Hyungkyu, et al.
Published: (2024)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
by: Cheng, Jianyi, et al.
Published: (2023)
by: Cheng, Jianyi, et al.
Published: (2023)
NPU Design for Diffusion Language Model Inference
by: Lou, Binglei, et al.
Published: (2026)
by: Lou, Binglei, et al.
Published: (2026)
Hardware Acceleration of Kolmogorov-Arnold Network (KAN) for Lightweight Edge Inference
by: Huang, Wei-Hsing, et al.
Published: (2024)
by: Huang, Wei-Hsing, et al.
Published: (2024)
High-Performance Data Mapping for BNNs on PCM-based Integrated Photonics
by: Shahroodi, Taha, et al.
Published: (2024)
by: Shahroodi, Taha, et al.
Published: (2024)
Online Training and Inference System on Edge FPGA Using Delayed Feedback Reservoir
by: Ikeda, Sosei, et al.
Published: (2025)
by: Ikeda, Sosei, et al.
Published: (2025)
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
by: Huang, Zhirui, et al.
Published: (2025)
by: Huang, Zhirui, et al.
Published: (2025)
VitaLLM: A Versatile and Tiny Accelerator for Mixed-Precision LLM Inference on Edge Devices
by: Lin, Zi-Wei, et al.
Published: (2026)
by: Lin, Zi-Wei, et al.
Published: (2026)
NVLLM: A 3D NAND-Centric Architecture Enabling Edge on-Device LLM Inference
by: Hao, Mingbo, et al.
Published: (2026)
by: Hao, Mingbo, et al.
Published: (2026)
CODO: An Automated Compiler for Comprehensive Dataflow Optimization
by: Zhang, Weichuang, et al.
Published: (2026)
by: Zhang, Weichuang, et al.
Published: (2026)
Compilation and Execution of an Embeddable YOLO-NAS on the VTA
by: Faure-Gignoux, Anthony, et al.
Published: (2026)
by: Faure-Gignoux, Anthony, et al.
Published: (2026)
Revet: A Language and Compiler for Dataflow Threads
by: Rucker, Alexander, et al.
Published: (2023)
by: Rucker, Alexander, et al.
Published: (2023)
Bombyx: OpenCilk Compilation for FPGA Hardware Acceleration
by: Shahawy, Mohamed, et al.
Published: (2025)
by: Shahawy, Mohamed, et al.
Published: (2025)
An FPGA Compiler for On-the-Fly Adaptive CNN Deployment and Reconfiguration
by: Mazouz, Alaa, et al.
Published: (2025)
by: Mazouz, Alaa, et al.
Published: (2025)
HillInfer: Efficient Long-Context LLM Inference on the Edge with Hierarchical KV Eviction using SmartSSD
by: Sun, He, et al.
Published: (2026)
by: Sun, He, et al.
Published: (2026)
In-Pipeline Integration of Digital In-Memory-Computing into RISC-V Vector Architecture to Accelerate Deep Learning
by: Spagnolo, Tommaso, et al.
Published: (2026)
by: Spagnolo, Tommaso, et al.
Published: (2026)
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
by: Sun, Xiaotian, et al.
Published: (2024)
by: Sun, Xiaotian, et al.
Published: (2024)
PD-Swap: Prefill-Decode Logic Swapping for End-to-End LLM Inference on Edge FPGAs via Dynamic Partial Reconfiguration
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
by: Yu, Jinxin, et al.
Published: (2026)
by: Yu, Jinxin, et al.
Published: (2026)
TriGen: NPU Architecture for End-to-End Acceleration of Large Language Models based on SW-HW Co-Design
by: Lee, Jonghun, et al.
Published: (2026)
by: Lee, Jonghun, et al.
Published: (2026)
Capstone: Power-Capped Pipelining for Coarse-Grained Reconfigurable Array Compilers
by: Yarzada, Sabrina, et al.
Published: (2026)
by: Yarzada, Sabrina, et al.
Published: (2026)
Building a Reusable and Extensible Automatic Compiler Infrastructure for Reconfigurable Devices
by: Zang, Zhenya, et al.
Published: (2023)
by: Zang, Zhenya, et al.
Published: (2023)
Implementing Keyword Spotting on the MCUX947 Microcontroller with Integrated NPU
by: Jakuš, Petar, et al.
Published: (2025)
by: Jakuš, Petar, et al.
Published: (2025)
Position Paper: From Edge AI to Adaptive Edge AI
by: Pittorino, Fabrizio, et al.
Published: (2026)
by: Pittorino, Fabrizio, et al.
Published: (2026)
Precision-Scalable Microscaling Datapaths with Optimized Reduction Tree for Efficient NPU Integration
by: Cuyckens, Stef, et al.
Published: (2025)
by: Cuyckens, Stef, et al.
Published: (2025)
OpenACM: An Open-Source SRAM-Based Approximate CiM Compiler
by: Zhou, Yiqi, et al.
Published: (2026)
by: Zhou, Yiqi, et al.
Published: (2026)
Similar Items
-
SILVIA: Automated Superword-Level Parallelism Exploitation via HLS-Specific LLVM Passes for Compute-Intensive FPGA Accelerators
by: Brignone, Giovanni, et al.
Published: (2024) -
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
by: Mueller, Lion, et al.
Published: (2025) -
A DSP shared is a DSP earned: HLS Task-Level Multi-Pumping for High-Performance Low-Resource Designs
by: Brignone, Giovanni, et al.
Published: (2023) -
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
by: Chen, Yuzong, et al.
Published: (2025) -
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
by: Heo, Guseul, et al.
Published: (2024)