Strix: Re-thinking NPU Reliability from a System Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918440908881920 |
|---|---|
| author | Guan, Jiapeng Zhang, Jie Zhou, Hao Wei, Ran You, Dean Wang, Hui Wang, Yingquan Wang, Tinglue Zhao, Xudong Li, Jing Jiang, Zhe |
| author_facet | Guan, Jiapeng Zhang, Jie Zhou, Hao Wei, Ran You, Dean Wang, Hui Wang, Yingquan Wang, Tinglue Zhao, Xudong Li, Jing Jiang, Zhe |
| contents | DNNs and LLMs increasingly rely on hardware accelerators, including in safety-critical domains, while technology scaling and growing model complexity make hardware faults more frequent. Existing system-level mechanisms typically treat the NPU as a monolithic unit, using coarse-grained replication that incurs prohibitive performance and hardware overheads, leaving a gap between reliability requirements and deployable solutions. To bridge this gap, we present Strix, a full-stack NPU reliability framework on an open-source SoC, spanning micro-architecture, ISA, and programming methods. Strix re-partitions the NPU along the system inference pipeline, identifies dominant failure modes, and attaches targeted safeguards, achieving sub-micro-second fault localisation, error detection, and correction with only 1.04$\times$ slowdown and minimal hardware overhead. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_10484 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Strix: Re-thinking NPU Reliability from a System Perspective Guan, Jiapeng Zhang, Jie Zhou, Hao Wei, Ran You, Dean Wang, Hui Wang, Yingquan Wang, Tinglue Zhao, Xudong Li, Jing Jiang, Zhe Hardware Architecture DNNs and LLMs increasingly rely on hardware accelerators, including in safety-critical domains, while technology scaling and growing model complexity make hardware faults more frequent. Existing system-level mechanisms typically treat the NPU as a monolithic unit, using coarse-grained replication that incurs prohibitive performance and hardware overheads, leaving a gap between reliability requirements and deployable solutions. To bridge this gap, we present Strix, a full-stack NPU reliability framework on an open-source SoC, spanning micro-architecture, ISA, and programming methods. Strix re-partitions the NPU along the system inference pipeline, identifies dominant failure modes, and attaches targeted safeguards, achieving sub-micro-second fault localisation, error detection, and correction with only 1.04$\times$ slowdown and minimal hardware overhead. |
| title | Strix: Re-thinking NPU Reliability from a System Perspective |
| topic | Hardware Architecture |
| url | https://arxiv.org/abs/2604.10484 |