On Fault Tolerance of Data Storage Systems: A Holistic Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Mai, Zhang, Duo, Dajani, Ahmed
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911039641092096
author Zheng, Mai
Zhang, Duo
Dajani, Ahmed
author_facet Zheng, Mai
Zhang, Duo
Dajani, Ahmed
contents Data storage systems serve as the foundation of digital society. The enormous data generated by people on a daily basis make the fault tolerance of data storage systems increasingly important. Unfortunately, modern storage systems consist of complicated hardware and software layers interacting with each other, which may contain latent bugs that elude extensive testing and lead to data corruption, system downtime, or even unrecoverable data loss in practice. In this chapter, we take a holistic view to introduce the typical architecture and major components of modern data storage systems (e.g., solid state drives, persistent memories, local file systems, and distributed storage management at scale). Next, we discuss a few representative bug detection and fault tolerance techniques across layers with a focus on issues that affect system recovery and data integrity. Finally, we conclude with open challenges and future work.
format Preprint
id arxiv_https___arxiv_org_abs_2507_03849
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On Fault Tolerance of Data Storage Systems: A Holistic Perspective
Zheng, Mai
Zhang, Duo
Dajani, Ahmed
Distributed, Parallel, and Cluster Computing
Data storage systems serve as the foundation of digital society. The enormous data generated by people on a daily basis make the fault tolerance of data storage systems increasingly important. Unfortunately, modern storage systems consist of complicated hardware and software layers interacting with each other, which may contain latent bugs that elude extensive testing and lead to data corruption, system downtime, or even unrecoverable data loss in practice. In this chapter, we take a holistic view to introduce the typical architecture and major components of modern data storage systems (e.g., solid state drives, persistent memories, local file systems, and distributed storage management at scale). Next, we discuss a few representative bug detection and fault tolerance techniques across layers with a focus on issues that affect system recovery and data integrity. Finally, we conclude with open challenges and future work.
title On Fault Tolerance of Data Storage Systems: A Holistic Perspective
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2507.03849