Investigating Memory Failure Prediction Across CPU Architectures

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yu, Qiao, Zhang, Wengui, Zhou, Min, Yu, Jialiang, Sheng, Zhenli, Bogatinovski, Jasmin, Cardoso, Jorge, Kao, Odej
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915063163518976
author Yu, Qiao
Zhang, Wengui
Zhou, Min
Yu, Jialiang
Sheng, Zhenli
Bogatinovski, Jasmin
Cardoso, Jorge
Kao, Odej
author_facet Yu, Qiao
Zhang, Wengui
Zhou, Min
Yu, Jialiang
Sheng, Zhenli
Bogatinovski, Jasmin
Cardoso, Jorge
Kao, Odej
contents Large-scale datacenters often experience memory failures, where Uncorrectable Errors (UEs) highlight critical malfunction in Dual Inline Memory Modules (DIMMs). Existing approaches primarily utilize Correctable Errors (CEs) to predict UEs, yet they typically neglect how these errors vary between different CPU architectures, especially in terms of Error Correction Code (ECC) applicability. In this paper, we investigate the correlation between CEs and UEs across different CPU architectures, including X86 and ARM. Our analysis identifies unique patterns of memory failure associated with each processor platform. Leveraging Machine Learning (ML) techniques on production datasets, we conduct the memory failure prediction in different processors' platforms, achieving up to 15% improvements in F1-score compared to the existing algorithm. Finally, an MLOps (Machine Learning Operations) framework is provided to consistently improve the failure prediction in the production environment.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05354
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Investigating Memory Failure Prediction Across CPU Architectures
Yu, Qiao
Zhang, Wengui
Zhou, Min
Yu, Jialiang
Sheng, Zhenli
Bogatinovski, Jasmin
Cardoso, Jorge
Kao, Odej
Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Large-scale datacenters often experience memory failures, where Uncorrectable Errors (UEs) highlight critical malfunction in Dual Inline Memory Modules (DIMMs). Existing approaches primarily utilize Correctable Errors (CEs) to predict UEs, yet they typically neglect how these errors vary between different CPU architectures, especially in terms of Error Correction Code (ECC) applicability. In this paper, we investigate the correlation between CEs and UEs across different CPU architectures, including X86 and ARM. Our analysis identifies unique patterns of memory failure associated with each processor platform. Leveraging Machine Learning (ML) techniques on production datasets, we conduct the memory failure prediction in different processors' platforms, achieving up to 15% improvements in F1-score compared to the existing algorithm. Finally, an MLOps (Machine Learning Operations) framework is provided to consistently improve the failure prediction in the production environment.
title Investigating Memory Failure Prediction Across CPU Architectures
topic Hardware Architecture
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2406.05354