The Energy-Throughput Trade-off in Lossless-Compressed Source Code Storage

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ferragina, Paolo, Tosoni, Francesco
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914265064013824
author Ferragina, Paolo
Tosoni, Francesco
author_facet Ferragina, Paolo
Tosoni, Francesco
contents Retrieving data from large-scale source code archives is vital for AI training, neural-based software analysis, and information retrieval, to cite a few. This paper studies and experiments with the design of a compressed key-value store for the indexing of large-scale source code datasets, evaluating its trade-off among three primary computational resources: (compressed) space occupancy, time, and energy efficiency. Extensive experiments on a national high-performance computing infrastructure demonstrate that different compression configurations yield distinct trade-offs, with high compression ratios and order-of-magnitude gains in retrieval throughput and energy efficiency. We also study data parallelism and show that, while it significantly improves speed, scaling energy efficiency is more difficult, reflecting the known non-energy-proportionality of modern hardware and challenging the assumption of a direct time-energy correlation. This work streamlines automation in energy-aware configuration tuning and standardized green benchmarking deployable in CI/CD pipelines, thus empowering system architects with a spectrum of Pareto-optimal energy-compression-throughput trade-offs and actionable guidelines for building sustainable, efficient storage backends for massive open-source code archival.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13220
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Energy-Throughput Trade-off in Lossless-Compressed Source Code Storage
Ferragina, Paolo
Tosoni, Francesco
Data Structures and Algorithms
Databases
Distributed, Parallel, and Cluster Computing
Performance
Software Engineering
E.2; H.3.2; H.3.3; D.2.8
Retrieving data from large-scale source code archives is vital for AI training, neural-based software analysis, and information retrieval, to cite a few. This paper studies and experiments with the design of a compressed key-value store for the indexing of large-scale source code datasets, evaluating its trade-off among three primary computational resources: (compressed) space occupancy, time, and energy efficiency. Extensive experiments on a national high-performance computing infrastructure demonstrate that different compression configurations yield distinct trade-offs, with high compression ratios and order-of-magnitude gains in retrieval throughput and energy efficiency. We also study data parallelism and show that, while it significantly improves speed, scaling energy efficiency is more difficult, reflecting the known non-energy-proportionality of modern hardware and challenging the assumption of a direct time-energy correlation. This work streamlines automation in energy-aware configuration tuning and standardized green benchmarking deployable in CI/CD pipelines, thus empowering system architects with a spectrum of Pareto-optimal energy-compression-throughput trade-offs and actionable guidelines for building sustainable, efficient storage backends for massive open-source code archival.
title The Energy-Throughput Trade-off in Lossless-Compressed Source Code Storage
topic Data Structures and Algorithms
Databases
Distributed, Parallel, and Cluster Computing
Performance
Software Engineering
E.2; H.3.2; H.3.3; D.2.8
url https://arxiv.org/abs/2601.13220