A Thorough Investigation of Content-Defined Chunking Algorithms for Data Deduplication

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gregoriadis, Marcel, Balduf, Leonhard, Scheuermann, Björn, Pouwelse, Johan
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916448044056576
author Gregoriadis, Marcel
Balduf, Leonhard
Scheuermann, Björn
Pouwelse, Johan
author_facet Gregoriadis, Marcel
Balduf, Leonhard
Scheuermann, Björn
Pouwelse, Johan
contents Data deduplication emerged as a powerful solution for reducing storage and bandwidth costs in cloud settings by eliminating redundancies at the level of chunks. This has spurred the development of numerous Content-Defined Chunking (CDC) algorithms over the past two decades. Despite advancements, the current state-of-the-art remains obscure, as a thorough and impartial analysis and comparison is lacking. We conduct a rigorous theoretical analysis and impartial experimental comparison of several leading CDC algorithms. Using four realistic datasets, we evaluate these algorithms against four key metrics: throughput, deduplication ratio, average chunk size, and chunk-size variance. Our analyses, in many instances, extend the findings of their original publications by reporting new results and putting existing ones into context. Moreover, we highlight limitations that have previously gone unnoticed. Our findings provide valuable insights that inform the selection and optimization of CDC algorithms for practical applications in data deduplication.
format Preprint
id arxiv_https___arxiv_org_abs_2409_06066
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Thorough Investigation of Content-Defined Chunking Algorithms for Data Deduplication
Gregoriadis, Marcel
Balduf, Leonhard
Scheuermann, Björn
Pouwelse, Johan
Distributed, Parallel, and Cluster Computing
Data deduplication emerged as a powerful solution for reducing storage and bandwidth costs in cloud settings by eliminating redundancies at the level of chunks. This has spurred the development of numerous Content-Defined Chunking (CDC) algorithms over the past two decades. Despite advancements, the current state-of-the-art remains obscure, as a thorough and impartial analysis and comparison is lacking. We conduct a rigorous theoretical analysis and impartial experimental comparison of several leading CDC algorithms. Using four realistic datasets, we evaluate these algorithms against four key metrics: throughput, deduplication ratio, average chunk size, and chunk-size variance. Our analyses, in many instances, extend the findings of their original publications by reporting new results and putting existing ones into context. Moreover, we highlight limitations that have previously gone unnoticed. Our findings provide valuable insights that inform the selection and optimization of CDC algorithms for practical applications in data deduplication.
title A Thorough Investigation of Content-Defined Chunking Algorithms for Data Deduplication
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2409.06066