From Characterization to Microarchitecture: Designing an Elegant and Reliable BFP-Based NPU
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jie, Guan, Jiapeng, Zhou, Hao, Han, Xiaomeng, Wang, Tinglue, Wei, Ran, Jiang, Zhe |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Strix: Re-thinking NPU Reliability from a System Perspective
by: Guan, Jiapeng, et al.
Published: (2026)
by: Guan, Jiapeng, et al.
Published: (2026)
Pushing the Limits of BFP on Narrow Precision LLM Inference
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
FlexStep: Enabling Flexible Error Detection in Multi/Many-core Real-time Systems
by: Wang, Tinglue, et al.
Published: (2025)
by: Wang, Tinglue, et al.
Published: (2025)
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
by: You, Dean, et al.
Published: (2025)
by: You, Dean, et al.
Published: (2025)
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
Benchmarking for Single Feature Attribution with Microarchitecture Cliffs
by: Zhen, Hao, et al.
Published: (2026)
by: Zhen, Hao, et al.
Published: (2026)
Automatic Microarchitecture-Aware Custom Instruction Design for RISC-V Processors
by: Rezunov, Evgenii, et al.
Published: (2025)
by: Rezunov, Evgenii, et al.
Published: (2025)
From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
by: Yu, Jinxin, et al.
Published: (2026)
by: Yu, Jinxin, et al.
Published: (2026)
FireGuard: A Generalized Microarchitecture for Fine-Grained Security Analysis on OoO Superscalar Cores
by: Jiang, Zhe, et al.
Published: (2025)
by: Jiang, Zhe, et al.
Published: (2025)
PyPIM: Integrating Digital Processing-in-Memory from Microarchitectural Design to Python Tensors
by: Leitersdorf, Orian, et al.
Published: (2023)
by: Leitersdorf, Orian, et al.
Published: (2023)
NeuroAI Temporal Neural Networks (NeuTNNs): Microarchitecture and Design Framework for Specialized Neuromorphic Processing Units
by: Venkatachalam, Shanmuga, et al.
Published: (2026)
by: Venkatachalam, Shanmuga, et al.
Published: (2026)
NVR: Vector Runahead on NPUs for Sparse Memory Access
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
Microarchitectural Co-Optimization for Sustained Throughput of RISC-V Multi-Lane Chaining Vector Processors
by: Wang, Weiying, et al.
Published: (2026)
by: Wang, Weiying, et al.
Published: (2026)
Rethinking Compute Substrates for 3D-Stacked Near-Memory LLM Decoding: Microarchitecture-Scheduling Co-Design
by: Ai, Chenyang, et al.
Published: (2026)
by: Ai, Chenyang, et al.
Published: (2026)
BBAL: A Bidirectional Block Floating Point-Based Quantisation Accelerator for Large Language Models
by: Han, Xiaomeng, et al.
Published: (2025)
by: Han, Xiaomeng, et al.
Published: (2025)
EONSim: An NPU Simulator for On-Chip Memory and Embedding Vector Operations
by: Choi, Sangun, et al.
Published: (2025)
by: Choi, Sangun, et al.
Published: (2025)
SemanticBBV: A Semantic Signature for Cross-Program Knowledge Reuse in Microarchitecture Simulation
by: Liu, Zhenguo, et al.
Published: (2025)
by: Liu, Zhenguo, et al.
Published: (2025)
HiMA: Hierarchical Quantum Microarchitecture for Qubit-Scaling and Quantum Process-Level Parallelism
by: Zhou, Qi, et al.
Published: (2024)
by: Zhou, Qi, et al.
Published: (2024)
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
by: Heo, Guseul, et al.
Published: (2024)
by: Heo, Guseul, et al.
Published: (2024)
IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
by: Seo, Minseok, et al.
Published: (2024)
by: Seo, Minseok, et al.
Published: (2024)
Regular-Dead on Arrival: Characterizing and Protecting Against Dead-Entry TLB Misses in GPU Microarchitectures
by: Anik, Shafayat Mowla, et al.
Published: (2026)
by: Anik, Shafayat Mowla, et al.
Published: (2026)
Insights from Rights and Wrongs: A Large Language Model for Solving Assertion Failures in RTL Design
by: Zhou, Jie, et al.
Published: (2025)
by: Zhou, Jie, et al.
Published: (2025)
MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs
by: Guan, Jiapeng, et al.
Published: (2024)
by: Guan, Jiapeng, et al.
Published: (2024)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
by: Zhou, Zhe, et al.
Published: (2024)
by: Zhou, Zhe, et al.
Published: (2024)
Evaluating the Effectiveness of Microarchitectural Hardware Fault Detection for Application-Specific Requirements
by: Papadopoulos, Konstantinos-Nikolaos, et al.
Published: (2024)
by: Papadopoulos, Konstantinos-Nikolaos, et al.
Published: (2024)
PacQ: A SIMT Microarchitecture for Efficient Dataflow in Hyper-asymmetric GEMMs
by: Yin, Ruokai, et al.
Published: (2025)
by: Yin, Ruokai, et al.
Published: (2025)
Microarchitecture Design and Benchmarking of Custom SHA-3 Instruction for RISC-V
by: Bolat, Alperen, et al.
Published: (2025)
by: Bolat, Alperen, et al.
Published: (2025)
ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
by: Ham, Hyungkyu, et al.
Published: (2024)
by: Ham, Hyungkyu, et al.
Published: (2024)
Supporting Secured Integration of Microarchitectural Defenses
by: Ramkrishnan, Kartik, et al.
Published: (2026)
by: Ramkrishnan, Kartik, et al.
Published: (2026)
A4: Microarchitecture-Aware LLC Management for Datacenter Servers with Emerging I/O Devices
by: Park, Haneul, et al.
Published: (2025)
by: Park, Haneul, et al.
Published: (2025)
Fine Grain 3D Integration for Microarchitecture Design Through Cube Packing Exploration
by: Liu, Yongxiang, et al.
Published: (2025)
by: Liu, Yongxiang, et al.
Published: (2025)
The Quest for Reliable AI Accelerators: Cross-Layer Evaluation and Design Optimization
by: Li, Meng, et al.
Published: (2026)
by: Li, Meng, et al.
Published: (2026)
Oreo: Protecting ASLR Against Microarchitectural Attacks (Extended Version)
by: Song, Shixin, et al.
Published: (2024)
by: Song, Shixin, et al.
Published: (2024)
Large Language Model for Verilog Generation with Code-Structure-Guided Reinforcement Learning
by: Wang, Ning, et al.
Published: (2024)
by: Wang, Ning, et al.
Published: (2024)
CXL-Interference: Analysis and Characterization in Modern Computer Systems
by: Mao, Shunyu, et al.
Published: (2024)
by: Mao, Shunyu, et al.
Published: (2024)
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
by: Wang, Zhao, et al.
Published: (2021)
by: Wang, Zhao, et al.
Published: (2021)
Rigorous Evaluation of Microarchitectural Side-Channels with Statistical Model Checking
by: Li, Weihang, et al.
Published: (2025)
by: Li, Weihang, et al.
Published: (2025)
FHECore: Rethinking GPU Microarchitecture for Fully Homomorphic Encryption
by: Daksha, Lohit, et al.
Published: (2026)
by: Daksha, Lohit, et al.
Published: (2026)
Tao: Re-Thinking DL-based Microarchitecture Simulation
by: Pandey, Santosh, et al.
Published: (2024)
by: Pandey, Santosh, et al.
Published: (2024)
TriGen: NPU Architecture for End-to-End Acceleration of Large Language Models based on SW-HW Co-Design
by: Lee, Jonghun, et al.
Published: (2026)
by: Lee, Jonghun, et al.
Published: (2026)
Similar Items
-
Strix: Re-thinking NPU Reliability from a System Perspective
by: Guan, Jiapeng, et al.
Published: (2026) -
Pushing the Limits of BFP on Narrow Precision LLM Inference
by: Wang, Hui, et al.
Published: (2025) -
FlexStep: Enabling Flexible Error Detection in Multi/Many-core Real-time Systems
by: Wang, Tinglue, et al.
Published: (2025) -
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
by: You, Dean, et al.
Published: (2025) -
Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
by: Wang, Xinyu, et al.
Published: (2026)