RedMulE-FT: A Reconfigurable Fault-Tolerant Matrix Multiplication Engine

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wiese, Philip, Item, Maurus, Bertaccini, Luca, Tortorella, Yvan, Garofalo, Angelo, Benini, Luca
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912337505550336
author Wiese, Philip
Item, Maurus
Bertaccini, Luca
Tortorella, Yvan
Garofalo, Angelo
Benini, Luca
author_facet Wiese, Philip
Item, Maurus
Bertaccini, Luca
Tortorella, Yvan
Garofalo, Angelo
Benini, Luca
contents As safety-critical applications increasingly rely on data-parallel floating-point computations, there is an increasing need for flexible and configurable fault tolerance in parallel floating-point accelerators such as tensor engines. While replication-based methods ensure reliability but incur high area and power costs, error correction codes lack the flexibility to trade off robustness against performance. This work presents RedMulE-FT, a runtime-configurable fault-tolerant extension of the RedMulE matrix multiplication accelerator, balancing fault tolerance, area overhead, and performance impacts. The fault tolerance mode is configured in a shadowed context register file before task execution. By combining replication with error-detecting codes to protect the data path, RedMulE-FT achieves an 11x uncorrected fault reduction with only 2.3% area overhead. Full protection extends to control signals, resulting in no functional errors after 1M injections during our extensive fault injection simulation campaign, with a total area overhead of 25.2% while maintaining a 500 MHz frequency in a 12 nm technology.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14399
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RedMulE-FT: A Reconfigurable Fault-Tolerant Matrix Multiplication Engine
Wiese, Philip
Item, Maurus
Bertaccini, Luca
Tortorella, Yvan
Garofalo, Angelo
Benini, Luca
Hardware Architecture
B.7.3; B.8.1
As safety-critical applications increasingly rely on data-parallel floating-point computations, there is an increasing need for flexible and configurable fault tolerance in parallel floating-point accelerators such as tensor engines. While replication-based methods ensure reliability but incur high area and power costs, error correction codes lack the flexibility to trade off robustness against performance. This work presents RedMulE-FT, a runtime-configurable fault-tolerant extension of the RedMulE matrix multiplication accelerator, balancing fault tolerance, area overhead, and performance impacts. The fault tolerance mode is configured in a shadowed context register file before task execution. By combining replication with error-detecting codes to protect the data path, RedMulE-FT achieves an 11x uncorrected fault reduction with only 2.3% area overhead. Full protection extends to control signals, resulting in no functional errors after 1M injections during our extensive fault injection simulation campaign, with a total area overhead of 25.2% while maintaining a 500 MHz frequency in a 12 nm technology.
title RedMulE-FT: A Reconfigurable Fault-Tolerant Matrix Multiplication Engine
topic Hardware Architecture
B.7.3; B.8.1
url https://arxiv.org/abs/2504.14399