FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dai, Huangliang, Wu, Shixun, Huang, Jiajun, Jian, Zizhe, Zhu, Yue, Hu, Haiyang, Chen, Zizhong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916894780424192
author Dai, Huangliang
Wu, Shixun
Huang, Jiajun
Jian, Zizhe
Zhu, Yue
Hu, Haiyang
Chen, Zizhong
author_facet Dai, Huangliang
Wu, Shixun
Huang, Jiajun
Jian, Zizhe
Zhu, Yue
Hu, Haiyang
Chen, Zizhong
contents Transformer models rely on High-Performance Computing (HPC) resources for inference, where soft errors are inevitable in large-scale systems, making the reliability of the model particularly critical. Existing fault tolerance frameworks for Transformers are designed at the operation level without architectural optimization, leading to significant computational and memory overhead, which in turn reduces protection efficiency and limits scalability to larger models. In this paper, we implement module-level protection for Transformers by treating the operations within the attention module as a single kernel and applying end-to-end fault tolerance. This method provides unified protection across multi-step computations, while achieving comprehensive coverage of potential errors in the nonlinear computations. For linear modules, we design a strided algorithm-based fault tolerance (ABFT) that avoids inter-thread communication. Experimental results show that our end-to-end fault tolerance achieves up to 7.56x speedup over traditional methods with an average fault tolerance overhead of 13.9%.
format Preprint
id arxiv_https___arxiv_org_abs_2504_02211
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant Attention
Dai, Huangliang
Wu, Shixun
Huang, Jiajun
Jian, Zizhe
Zhu, Yue
Hu, Haiyang
Chen, Zizhong
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
Transformer models rely on High-Performance Computing (HPC) resources for inference, where soft errors are inevitable in large-scale systems, making the reliability of the model particularly critical. Existing fault tolerance frameworks for Transformers are designed at the operation level without architectural optimization, leading to significant computational and memory overhead, which in turn reduces protection efficiency and limits scalability to larger models. In this paper, we implement module-level protection for Transformers by treating the operations within the attention module as a single kernel and applying end-to-end fault tolerance. This method provides unified protection across multi-step computations, while achieving comprehensive coverage of potential errors in the nonlinear computations. For linear modules, we design a strided algorithm-based fault tolerance (ABFT) that avoids inter-thread communication. Experimental results show that our end-to-end fault tolerance achieves up to 7.56x speedup over traditional methods with an average fault tolerance overhead of 13.9%.
title FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant Attention
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2504.02211