Precision-Scalable Microscaling Datapaths with Optimized Reduction Tree for Efficient NPU Integration
Fuente:
arXiv
Saved in:
| Main Authors: | Cuyckens, Stef, Yi, Xiaoling, Geens, Robin, Dumoulin, Joren, Wiesner, Martin, Fang, Chao, Verhelst, Marian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Precision-Scalable Hardware for Microscaling (MX) Processing in Robotics Learning
by: Cuyckens, Stef, et al.
Published: (2025)
by: Cuyckens, Stef, et al.
Published: (2025)
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference
by: Geens, Robin, et al.
Published: (2026)
by: Geens, Robin, et al.
Published: (2026)
Hardware-Centric Analysis of DeepSeek's Multi-Head Latent Attention
by: Geens, Robin, et al.
Published: (2025)
by: Geens, Robin, et al.
Published: (2025)
iEEG Seizure Detection with a Sparse Hyperdimensional Computing Accelerator
by: Cuyckens, Stef, et al.
Published: (2025)
by: Cuyckens, Stef, et al.
Published: (2025)
Fine-Grained Fusion: The Missing Piece in Area-Efficient State Space Model Acceleration
by: Geens, Robin, et al.
Published: (2025)
by: Geens, Robin, et al.
Published: (2025)
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
by: Dumoulin, Joren, et al.
Published: (2025)
by: Dumoulin, Joren, et al.
Published: (2025)
A 16 nm 1.60TOPS/W High Utilization DNN Accelerator with 3D Spatial Data Reuse and Efficient Shared Memory Access
by: Yi, Xiaoling, et al.
Published: (2026)
by: Yi, Xiaoling, et al.
Published: (2026)
Refining Datapath for Microscaling ViTs
by: Xiao, Can, et al.
Published: (2025)
by: Xiao, Can, et al.
Published: (2025)
The Hyperscale Lottery: How State-Space Models Have Sacrificed Edge Efficiency
by: Geens, Robin, et al.
Published: (2026)
by: Geens, Robin, et al.
Published: (2026)
An Open-Source HW-SW Co-Development Framework Enabling Efficient Multi-Accelerator Systems
by: Antonio, Ryan Albert, et al.
Published: (2025)
by: Antonio, Ryan Albert, et al.
Published: (2025)
OpenGeMM: A High-Utilization GeMM Accelerator Generator with Lightweight RISC-V Control and Tight Memory Coupling
by: Yi, Xiaoling, et al.
Published: (2024)
by: Yi, Xiaoling, et al.
Published: (2024)
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
by: Fang, Chao, et al.
Published: (2024)
by: Fang, Chao, et al.
Published: (2024)
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
by: Chen, Yuzong, et al.
Published: (2025)
by: Chen, Yuzong, et al.
Published: (2025)
Analog or Digital In-memory Computing? Benchmarking through Quantitative Modeling
by: Sun, Jiacong, et al.
Published: (2024)
by: Sun, Jiacong, et al.
Published: (2024)
DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow Accelerators
by: Yi, Xiaoling, et al.
Published: (2025)
by: Yi, Xiaoling, et al.
Published: (2025)
Combining Power and Arithmetic Optimization via Datapath Rewriting
by: Coward, Samuel, et al.
Published: (2024)
by: Coward, Samuel, et al.
Published: (2024)
Torrent: A Distributed DMA for Efficient and Flexible Point-to-Multipoint Data Movement
by: Deng, Yunhao, et al.
Published: (2025)
by: Deng, Yunhao, et al.
Published: (2025)
How to keep pushing ML accelerator performance? Know your rooflines!
by: Verhelst, Marian, et al.
Published: (2025)
by: Verhelst, Marian, et al.
Published: (2025)
Scalable Wavelength Arbitration for Microring-based DWDM Transceivers
by: Choi, Sunjin, et al.
Published: (2024)
by: Choi, Sunjin, et al.
Published: (2024)
MEDUSA: Scalable Biometric Sensing in the Wild through Distributed MIMO Radars
by: Li, Yilong, et al.
Published: (2023)
by: Li, Yilong, et al.
Published: (2023)
Optimizing Layer-Fused Scheduling of Transformer Networks on Multi-accelerator Platforms
by: Colleman, Steven, et al.
Published: (2024)
by: Colleman, Steven, et al.
Published: (2024)
Low-latency D-MIMO Localization using Distributed Scalable Message-Passing Algorithm
by: Iancu, Dumitra, et al.
Published: (2025)
by: Iancu, Dumitra, et al.
Published: (2025)
A Scalable RISC-V Vector Processor Enabling Efficient Multi-Precision DNN Inference
by: Wang, Chuanning, et al.
Published: (2024)
by: Wang, Chuanning, et al.
Published: (2024)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
by: Kong, Fanchen, et al.
Published: (2025)
by: Kong, Fanchen, et al.
Published: (2025)
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
by: Cheng, Jianyi, et al.
Published: (2023)
by: Cheng, Jianyi, et al.
Published: (2023)
IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
by: Seo, Minseok, et al.
Published: (2024)
by: Seo, Minseok, et al.
Published: (2024)
SPEED: A Scalable RISC-V Vector Processor Enabling Efficient Multi-Precision DNN Inference
by: Wang, Chuanning, et al.
Published: (2024)
by: Wang, Chuanning, et al.
Published: (2024)
Design and In-training Optimization of Binary Search ADC for Flexible Classifiers
by: Duarte, Paula Carolina Lozano, et al.
Published: (2024)
by: Duarte, Paula Carolina Lozano, et al.
Published: (2024)
Implementation and Optimization of HQC Decoding on NPU-Integrated Devices
by: Chau, Vu Minh, et al.
Published: (2026)
by: Chau, Vu Minh, et al.
Published: (2026)
Pack my weights and run! Minimizing overheads for in-memory computing accelerators
by: Houshmand, Pouya, et al.
Published: (2024)
by: Houshmand, Pouya, et al.
Published: (2024)
Slimmed optical neural networks with multiplexed neuron sets and a corresponding backpropagation training algorithm
by: Liu, Yi-Feng, et al.
Published: (2023)
by: Liu, Yi-Feng, et al.
Published: (2023)
Bio-RV: Low-Power Resource-Efficient RISC-V Processor for Biomedical Applications
by: Sharma, Vijay Pratap, et al.
Published: (2026)
by: Sharma, Vijay Pratap, et al.
Published: (2026)
MUSIC-lite: Efficient MUSIC using Approximate Computing: An OFDM Radar Case Study
by: Bhattacharjya, Rajat, et al.
Published: (2024)
by: Bhattacharjya, Rajat, et al.
Published: (2024)
FieldHAR: A Fully Integrated End-to-end RTL Framework for Human Activity Recognition with Neural Networks from Heterogeneous Sensors
by: Liu, Mengxi, et al.
Published: (2023)
by: Liu, Mengxi, et al.
Published: (2023)
ONE-SA: Enabling Nonlinear Operations in Systolic Arrays for Efficient and Flexible Neural Network Inference
by: Sun, Ruiqi, et al.
Published: (2024)
by: Sun, Ruiqi, et al.
Published: (2024)
Variable Point: A Number Format for Area- and Energy-Efficient Multiplication of High-Dynamic-Range Numbers
by: Mirfarshbafan, Seyed Hadi, et al.
Published: (2025)
by: Mirfarshbafan, Seyed Hadi, et al.
Published: (2025)
VMXDOTP: A RISC-V Vector ISA Extension for Efficient Microscaling (MX) Format Acceleration
by: Wipfli, Max, et al.
Published: (2026)
by: Wipfli, Max, et al.
Published: (2026)
A 0.03${mm}^2$ 100-250MHz Charge-Pump or Amplifier-Less Integrating Sub-Sampling PLL for Ultra-low Power Communication and Computing
by: Ray, Yudhajit, et al.
Published: (2024)
by: Ray, Yudhajit, et al.
Published: (2024)
PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage Fusion
by: Wang, Huizheng, et al.
Published: (2025)
by: Wang, Huizheng, et al.
Published: (2025)
M2XFP: A Metadata-Augmented Microscaling Data Format for Efficient Low-bit Quantization
by: Hu, Weiming, et al.
Published: (2026)
by: Hu, Weiming, et al.
Published: (2026)
Similar Items
-
Efficient Precision-Scalable Hardware for Microscaling (MX) Processing in Robotics Learning
by: Cuyckens, Stef, et al.
Published: (2025) -
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference
by: Geens, Robin, et al.
Published: (2026) -
Hardware-Centric Analysis of DeepSeek's Multi-Head Latent Attention
by: Geens, Robin, et al.
Published: (2025) -
iEEG Seizure Detection with a Sparse Hyperdimensional Computing Accelerator
by: Cuyckens, Stef, et al.
Published: (2025) -
Fine-Grained Fusion: The Missing Piece in Area-Efficient State Space Model Acceleration
by: Geens, Robin, et al.
Published: (2025)