LOOPerSet: A Large-Scale Dataset for Data-Driven Polyhedral Compiler Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Merouani, Massinissa, Boudaoud, Afif, Baghdadi, Riyadh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912790661300224
author Merouani, Massinissa
Boudaoud, Afif
Baghdadi, Riyadh
author_facet Merouani, Massinissa
Boudaoud, Afif
Baghdadi, Riyadh
contents The advancement of machine learning for compiler optimization, particularly within the polyhedral model, is constrained by the scarcity of large-scale, public performance datasets. This data bottleneck forces researchers to undertake costly data generation campaigns, slowing down innovation and hindering reproducible research learned code optimization. To address this gap, we introduce LOOPerSet, a new public dataset containing 28 million labeled data points derived from 220,000 unique, synthetically generated polyhedral programs. Each data point maps a program and a complex sequence of semantics-preserving transformations (such as fusion, skewing, tiling, and parallelism)to a ground truth performance measurement (execution time). The scale and diversity of LOOPerSet make it a valuable resource for training and evaluating learned cost models, benchmarking new model architectures, and exploring the frontiers of automated polyhedral scheduling. The dataset is released under a permissive license to foster reproducible research and lower the barrier to entry for data-driven compiler optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10209
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LOOPerSet: A Large-Scale Dataset for Data-Driven Polyhedral Compiler Optimization
Merouani, Massinissa
Boudaoud, Afif
Baghdadi, Riyadh
Programming Languages
Machine Learning
Performance
The advancement of machine learning for compiler optimization, particularly within the polyhedral model, is constrained by the scarcity of large-scale, public performance datasets. This data bottleneck forces researchers to undertake costly data generation campaigns, slowing down innovation and hindering reproducible research learned code optimization. To address this gap, we introduce LOOPerSet, a new public dataset containing 28 million labeled data points derived from 220,000 unique, synthetically generated polyhedral programs. Each data point maps a program and a complex sequence of semantics-preserving transformations (such as fusion, skewing, tiling, and parallelism)to a ground truth performance measurement (execution time). The scale and diversity of LOOPerSet make it a valuable resource for training and evaluating learned cost models, benchmarking new model architectures, and exploring the frontiers of automated polyhedral scheduling. The dataset is released under a permissive license to foster reproducible research and lower the barrier to entry for data-driven compiler optimization.
title LOOPerSet: A Large-Scale Dataset for Data-Driven Polyhedral Compiler Optimization
topic Programming Languages
Machine Learning
Performance
url https://arxiv.org/abs/2510.10209