Strix: Re-thinking NPU Reliability from a System Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Guan, Jiapeng, Zhang, Jie, Zhou, Hao, Wei, Ran, You, Dean, Wang, Hui, Wang, Yingquan, Wang, Tinglue, Zhao, Xudong, Li, Jing, Jiang, Zhe |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Characterization to Microarchitecture: Designing an Elegant and Reliable BFP-Based NPU
by: Zhang, Jie, et al.
Published: (2026)
by: Zhang, Jie, et al.
Published: (2026)
MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs
by: Guan, Jiapeng, et al.
Published: (2024)
by: Guan, Jiapeng, et al.
Published: (2024)
FlexStep: Enabling Flexible Error Detection in Multi/Many-core Real-time Systems
by: Wang, Tinglue, et al.
Published: (2025)
by: Wang, Tinglue, et al.
Published: (2025)
MEEK: Re-thinking Heterogeneous Parallel Error Detection Architecture for Real-World OoO Superscalar Processors
by: Jiang, Zhe, et al.
Published: (2025)
by: Jiang, Zhe, et al.
Published: (2025)
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
by: You, Dean, et al.
Published: (2025)
by: You, Dean, et al.
Published: (2025)
NVR: Vector Runahead on NPUs for Sparse Memory Access
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
MEIC: Re-thinking RTL Debug Automation using LLMs
by: Xu, Ke, et al.
Published: (2024)
by: Xu, Ke, et al.
Published: (2024)
ISAAC: Intelligent, Scalable, Agile, and Accelerated CPU Verification via LLM-aided FPGA Parallelism
by: Sun, Jialin, et al.
Published: (2025)
by: Sun, Jialin, et al.
Published: (2025)
IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
by: Seo, Minseok, et al.
Published: (2024)
by: Seo, Minseok, et al.
Published: (2024)
EONSim: An NPU Simulator for On-Chip Memory and Embedding Vector Operations
by: Choi, Sangun, et al.
Published: (2025)
by: Choi, Sangun, et al.
Published: (2025)
Re-thinking Memory-Bound Limitations in CGRAs
by: Liu, Xiangfeng, et al.
Published: (2025)
by: Liu, Xiangfeng, et al.
Published: (2025)
Large Language Model for Verilog Generation with Code-Structure-Guided Reinforcement Learning
by: Wang, Ning, et al.
Published: (2024)
by: Wang, Ning, et al.
Published: (2024)
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
by: Heo, Guseul, et al.
Published: (2024)
by: Heo, Guseul, et al.
Published: (2024)
ReChisel: Effective Automatic Chisel Code Generation by LLM with Reflection
by: Niu, Juxin, et al.
Published: (2025)
by: Niu, Juxin, et al.
Published: (2025)
FPGA-based Emulation and Device-Side Management for CXL-based Memory Tiering Systems
by: Chen, Yiqi, et al.
Published: (2025)
by: Chen, Yiqi, et al.
Published: (2025)
Insights from Verification: Training a Verilog Generation LLM with Reinforcement Learning with Testbench Feedback
by: Wang, Ning, et al.
Published: (2025)
by: Wang, Ning, et al.
Published: (2025)
Location is Key: Leveraging Large Language Model for Functional Bug Localization in Verilog
by: Yao, Bingkun, et al.
Published: (2024)
by: Yao, Bingkun, et al.
Published: (2024)
Insights from Rights and Wrongs: A Large Language Model for Solving Assertion Failures in RTL Design
by: Zhou, Jie, et al.
Published: (2025)
by: Zhou, Jie, et al.
Published: (2025)
ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
by: Ham, Hyungkyu, et al.
Published: (2024)
by: Ham, Hyungkyu, et al.
Published: (2024)
BBAL: A Bidirectional Block Floating Point-Based Quantisation Accelerator for Large Language Models
by: Han, Xiaomeng, et al.
Published: (2025)
by: Han, Xiaomeng, et al.
Published: (2025)
From Buffers to Registers: Unlocking Fine-Grained FlashAttention with Hybrid-Bonded 3D NPU Co-Design
by: Yu, Jinxin, et al.
Published: (2026)
by: Yu, Jinxin, et al.
Published: (2026)
RePart: Efficient Hypergraph Partitioning with Logic Replication Optimization for Multi-FPGA System
by: Fu, Zizhuo, et al.
Published: (2026)
by: Fu, Zizhuo, et al.
Published: (2026)
CREATE: Cross-Layer Resilience Characterization and Optimization for Efficient yet Reliable Embodied AI Systems
by: Xie, Tong, et al.
Published: (2026)
by: Xie, Tong, et al.
Published: (2026)
UVLLM: An Automated Universal RTL Verification Framework using LLMs
by: Hu, Yuchen, et al.
Published: (2024)
by: Hu, Yuchen, et al.
Published: (2024)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
by: Zhou, Zhe, et al.
Published: (2024)
by: Zhou, Zhe, et al.
Published: (2024)
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
by: Wang, Zhao, et al.
Published: (2021)
by: Wang, Zhao, et al.
Published: (2021)
Implementation and Optimization of HQC Decoding on NPU-Integrated Devices
by: Chau, Vu Minh, et al.
Published: (2026)
by: Chau, Vu Minh, et al.
Published: (2026)
Aging Aware Adaptive Voltage Scaling for Reliable and Efficient AI Accelerators
by: Xie, Tong, et al.
Published: (2026)
by: Xie, Tong, et al.
Published: (2026)
VeriDebug: A Unified LLM for Verilog Debugging via Contrastive Embedding and Guided Correction
by: Wang, Ning, et al.
Published: (2025)
by: Wang, Ning, et al.
Published: (2025)
The Quest for Reliable AI Accelerators: Cross-Layer Evaluation and Design Optimization
by: Li, Meng, et al.
Published: (2026)
by: Li, Meng, et al.
Published: (2026)
DRIFT: Harnessing Inherent Fault Tolerance for Efficient and Reliable Diffusion Model Inference
by: Wen, Jinqi, et al.
Published: (2026)
by: Wen, Jinqi, et al.
Published: (2026)
ReaLM: Reliable and Efficient Large Language Model Inference with Statistical Algorithm-Based Fault Tolerance
by: Xie, Tong, et al.
Published: (2025)
by: Xie, Tong, et al.
Published: (2025)
ReDas: A Lightweight Architecture for Supporting Fine-Grained Reshaping and Multiple Dataflows on Systolic Array
by: Han, Meng, et al.
Published: (2023)
by: Han, Meng, et al.
Published: (2023)
eIQ Neutron: Redefining Edge-AI Inference with Integrated NPU and Compiler Innovations
by: Bamberg, Lennart, et al.
Published: (2025)
by: Bamberg, Lennart, et al.
Published: (2025)
VerilogMonkey: Exploring Parallel Scaling for Automated Verilog Code Generation with LLMs
by: Niu, Juxin, et al.
Published: (2025)
by: Niu, Juxin, et al.
Published: (2025)
RARO: Reliability-aware Conversion with Enhanced Read Performance for QLC SSDs
by: Wang, Yanyun, et al.
Published: (2025)
by: Wang, Yanyun, et al.
Published: (2025)
UVMarvel: an Automated LLM-aided UVM Machine for Subsystem-level RTL Verification
by: Ye, Junhao, et al.
Published: (2026)
by: Ye, Junhao, et al.
Published: (2026)
CXL-Interference: Analysis and Characterization in Modern Computer Systems
by: Mao, Shunyu, et al.
Published: (2024)
by: Mao, Shunyu, et al.
Published: (2024)
Changing the Game: The Bounce-Bind Ising Machine
by: Zhang, Haiyang, et al.
Published: (2026)
by: Zhang, Haiyang, et al.
Published: (2026)
An Architectural Error Metric for CNN-Oriented Approximate Multipliers
by: Liu, Ao, et al.
Published: (2024)
by: Liu, Ao, et al.
Published: (2024)
Similar Items
-
From Characterization to Microarchitecture: Designing an Elegant and Reliable BFP-Based NPU
by: Zhang, Jie, et al.
Published: (2026) -
MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs
by: Guan, Jiapeng, et al.
Published: (2024) -
FlexStep: Enabling Flexible Error Detection in Multi/Many-core Real-time Systems
by: Wang, Tinglue, et al.
Published: (2025) -
MEEK: Re-thinking Heterogeneous Parallel Error Detection Architecture for Real-World OoO Superscalar Processors
by: Jiang, Zhe, et al.
Published: (2025) -
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
by: You, Dean, et al.
Published: (2025)