Strix: Re-thinking NPU Reliability from a System Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guan, Jiapeng, Zhang, Jie, Zhou, Hao, Wei, Ran, You, Dean, Wang, Hui, Wang, Yingquan, Wang, Tinglue, Zhao, Xudong, Li, Jing, Jiang, Zhe
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918440908881920
author Guan, Jiapeng
Zhang, Jie
Zhou, Hao
Wei, Ran
You, Dean
Wang, Hui
Wang, Yingquan
Wang, Tinglue
Zhao, Xudong
Li, Jing
Jiang, Zhe
author_facet Guan, Jiapeng
Zhang, Jie
Zhou, Hao
Wei, Ran
You, Dean
Wang, Hui
Wang, Yingquan
Wang, Tinglue
Zhao, Xudong
Li, Jing
Jiang, Zhe
contents DNNs and LLMs increasingly rely on hardware accelerators, including in safety-critical domains, while technology scaling and growing model complexity make hardware faults more frequent. Existing system-level mechanisms typically treat the NPU as a monolithic unit, using coarse-grained replication that incurs prohibitive performance and hardware overheads, leaving a gap between reliability requirements and deployable solutions. To bridge this gap, we present Strix, a full-stack NPU reliability framework on an open-source SoC, spanning micro-architecture, ISA, and programming methods. Strix re-partitions the NPU along the system inference pipeline, identifies dominant failure modes, and attaches targeted safeguards, achieving sub-micro-second fault localisation, error detection, and correction with only 1.04$\times$ slowdown and minimal hardware overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10484
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Strix: Re-thinking NPU Reliability from a System Perspective
Guan, Jiapeng
Zhang, Jie
Zhou, Hao
Wei, Ran
You, Dean
Wang, Hui
Wang, Yingquan
Wang, Tinglue
Zhao, Xudong
Li, Jing
Jiang, Zhe
Hardware Architecture
DNNs and LLMs increasingly rely on hardware accelerators, including in safety-critical domains, while technology scaling and growing model complexity make hardware faults more frequent. Existing system-level mechanisms typically treat the NPU as a monolithic unit, using coarse-grained replication that incurs prohibitive performance and hardware overheads, leaving a gap between reliability requirements and deployable solutions. To bridge this gap, we present Strix, a full-stack NPU reliability framework on an open-source SoC, spanning micro-architecture, ISA, and programming methods. Strix re-partitions the NPU along the system inference pipeline, identifies dominant failure modes, and attaches targeted safeguards, achieving sub-micro-second fault localisation, error detection, and correction with only 1.04$\times$ slowdown and minimal hardware overhead.
title Strix: Re-thinking NPU Reliability from a System Perspective
topic Hardware Architecture
url https://arxiv.org/abs/2604.10484