Benchmarking Ultra-Low-Power $μ$NPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Millar, Josh, Huang, Yushan, Sethi, Sarab, Haddadi, Hamed, Madhavapeddy, Anil |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Energy-Aware Deep Learning on Resource-Constrained Hardware
by: Millar, Josh, et al.
Published: (2025)
by: Millar, Josh, et al.
Published: (2025)
Terracorder: Sense Long and Prosper
by: Millar, Josh, et al.
Published: (2024)
by: Millar, Josh, et al.
Published: (2024)
Low-Energy On-Device Personalization for MCUs
by: Huang, Yushan, et al.
Published: (2024)
by: Huang, Yushan, et al.
Published: (2024)
From Principles to Practice: A Systematic Study of LLM Serving on Multi-core NPUs
by: Zhu, Tianhao, et al.
Published: (2025)
by: Zhu, Tianhao, et al.
Published: (2025)
SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR
by: Wang, Erwei, et al.
Published: (2025)
by: Wang, Erwei, et al.
Published: (2025)
Exploration of Low-Power Flexible Stress Monitoring Classifiers for Conformal Wearables
by: Afentaki, Florentia, et al.
Published: (2025)
by: Afentaki, Florentia, et al.
Published: (2025)
Low Power Vision Transformer Accelerator with Hardware-Aware Pruning and Optimized Dataflow
by: Hsiung, Ching-Lin, et al.
Published: (2025)
by: Hsiung, Ching-Lin, et al.
Published: (2025)
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
by: Andronic, Marta, et al.
Published: (2023)
by: Andronic, Marta, et al.
Published: (2023)
Improving the Performance and Learning Stability of Parallelizable RNNs Designed for Ultra-Low Power Applications
by: Brandoit, Julien, et al.
Published: (2026)
by: Brandoit, Julien, et al.
Published: (2026)
Towards Real-Time ECG and EMG Modeling on $μ$NPUs
by: Millar, Josh, et al.
Published: (2026)
by: Millar, Josh, et al.
Published: (2026)
Decentor-V: Lightweight ML Training on Low-Power RISC-V Edge Devices
by: Ribeiro, Marcelo, et al.
Published: (2025)
by: Ribeiro, Marcelo, et al.
Published: (2025)
PowerGenie: Analytically-Guided Evolutionary Discovery of Superior Reconfigurable Power Converters
by: Gao, Jian, et al.
Published: (2026)
by: Gao, Jian, et al.
Published: (2026)
ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
by: Yang, Jinwu, et al.
Published: (2026)
by: Yang, Jinwu, et al.
Published: (2026)
Layer-wise Weight Selection for Power-Efficient Neural Network Acceleration
by: Fang, Jiaxun, et al.
Published: (2025)
by: Fang, Jiaxun, et al.
Published: (2025)
SONIQ: System-Optimized Noise-Injected Ultra-Low-Precision Quantization with Full-Precision Parity
by: Zhou, Cyrus, et al.
Published: (2023)
by: Zhou, Cyrus, et al.
Published: (2023)
RealBench: Benchmarking Verilog Generation Models with Real-World IP Designs
by: Jin, Pengwei, et al.
Published: (2025)
by: Jin, Pengwei, et al.
Published: (2025)
M100: An Orchestrated Dataflow Architecture Powering General AI Computing
by: Xie, Yan, et al.
Published: (2026)
by: Xie, Yan, et al.
Published: (2026)
PICBench: Benchmarking LLMs for Photonic Integrated Circuits Design
by: Wu, Yuchao, et al.
Published: (2025)
by: Wu, Yuchao, et al.
Published: (2025)
T-MAN: Enabling End-to-End Low-Bit LLM Inference on NPUs via Unified Table Lookup
by: Wei, Jianyu, et al.
Published: (2025)
by: Wei, Jianyu, et al.
Published: (2025)
PolyLUT: Ultra-low Latency Polynomial Inference with Hardware-Aware Structured Pruning
by: Andronic, Marta, et al.
Published: (2025)
by: Andronic, Marta, et al.
Published: (2025)
AnalogNAS-Bench: A NAS Benchmark for Analog In-Memory Computing
by: Bessalah, Aniss, et al.
Published: (2025)
by: Bessalah, Aniss, et al.
Published: (2025)
μRL: Discovering Transient Execution Vulnerabilities Using Reinforcement Learning
by: Tol, M. Caner, et al.
Published: (2025)
by: Tol, M. Caner, et al.
Published: (2025)
RTL-Repo: A Benchmark for Evaluating LLMs on Large-Scale RTL Design Projects
by: Allam, Ahmed, et al.
Published: (2024)
by: Allam, Ahmed, et al.
Published: (2024)
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
by: Alexandridis, Kosmas, et al.
Published: (2025)
by: Alexandridis, Kosmas, et al.
Published: (2025)
EvolveGen: Algorithmic Level Hardware Model Checking Benchmark Generation through Reinforcement Learning
by: Hu, Guangyu, et al.
Published: (2026)
by: Hu, Guangyu, et al.
Published: (2026)
TRAM: Training Approximate Multiplier Structures for Low-Power AI Accelerators
by: Meng, Chang, et al.
Published: (2026)
by: Meng, Chang, et al.
Published: (2026)
Efficient FPGA Implementation of Time-Domain Popcount for Low-Complexity Machine Learning
by: Duan, Shengyu, et al.
Published: (2025)
by: Duan, Shengyu, et al.
Published: (2025)
MGS: Markov Greedy Sums for Accurate Low-Bitwidth Floating-Point Accumulation
by: Natesh, Vikas, et al.
Published: (2025)
by: Natesh, Vikas, et al.
Published: (2025)
GCN-ABFT: Low-Cost Online Error Checking for Graph Convolutional Networks
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
LLM-VeriPPA: Power, Performance, and Area Optimization aware Verilog Code Generation with Large Language Models
by: Thorat, Kiran, et al.
Published: (2025)
by: Thorat, Kiran, et al.
Published: (2025)
VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers
by: Wang, Run, et al.
Published: (2025)
by: Wang, Run, et al.
Published: (2025)
HW-SW Optimization of DNNs for Privacy-preserving People Counting on Low-resolution Infrared Arrays
by: Risso, Matteo, et al.
Published: (2024)
by: Risso, Matteo, et al.
Published: (2024)
Architectural Implications of Neural Network Inference for High Data-Rate, Low-Latency Scientific Applications
by: Weng, Olivia, et al.
Published: (2024)
by: Weng, Olivia, et al.
Published: (2024)
MEADOW: Memory-efficient Dataflow and Data Packing for Low Power Edge LLMs
by: Moitra, Abhishek, et al.
Published: (2025)
by: Moitra, Abhishek, et al.
Published: (2025)
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
by: Xie, Xilong, et al.
Published: (2025)
by: Xie, Xilong, et al.
Published: (2025)
Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification
by: Pinckney, Nathaniel, et al.
Published: (2025)
by: Pinckney, Nathaniel, et al.
Published: (2025)
SeVeDo: A Heterogeneous Transformer Accelerator for Low-Bit Inference via Hierarchical Group Quantization and SVD-Guided Mixed Precision
by: Choi, Yuseon, et al.
Published: (2025)
by: Choi, Yuseon, et al.
Published: (2025)
Striking the Balance: GEMM Performance Optimization Across Generations of Ryzen AI NPUs
by: Taka, Endri, et al.
Published: (2025)
by: Taka, Endri, et al.
Published: (2025)
MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs
by: Wu, Haoran, et al.
Published: (2026)
by: Wu, Haoran, et al.
Published: (2026)
Similar Items
-
Energy-Aware Deep Learning on Resource-Constrained Hardware
by: Millar, Josh, et al.
Published: (2025) -
Terracorder: Sense Long and Prosper
by: Millar, Josh, et al.
Published: (2024) -
Low-Energy On-Device Personalization for MCUs
by: Huang, Yushan, et al.
Published: (2024) -
From Principles to Practice: A Systematic Study of LLM Serving on Multi-core NPUs
by: Zhu, Tianhao, et al.
Published: (2025) -
SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
by: Zhang, Jintao, et al.
Published: (2026)