PBench: Workload Synthesizer with Real Statistics for Cloud Analytics Benchmarking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yan, Liu, Chunwei, Urgaonkar, Bhuvan, Wang, Zhengle, Mueller, Magnus, Zhang, Chao, Zhang, Songyue, Pfeil, Pascal, Horn, Dominik, Liu, Zhengchun, Pagano, Davide, Kraska, Tim, Madden, Samuel, Fan, Ju
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909653939519488
author Zhou, Yan
Liu, Chunwei
Urgaonkar, Bhuvan
Wang, Zhengle
Mueller, Magnus
Zhang, Chao
Zhang, Songyue
Pfeil, Pascal
Horn, Dominik
Liu, Zhengchun
Pagano, Davide
Kraska, Tim
Madden, Samuel
Fan, Ju
author_facet Zhou, Yan
Liu, Chunwei
Urgaonkar, Bhuvan
Wang, Zhengle
Mueller, Magnus
Zhang, Chao
Zhang, Songyue
Pfeil, Pascal
Horn, Dominik
Liu, Zhengchun
Pagano, Davide
Kraska, Tim
Madden, Samuel
Fan, Ju
contents Cloud service providers commonly use standard benchmarks like TPC-H and TPC-DS to evaluate and optimize cloud data analytics systems. However, these benchmarks rely on fixed query patterns and fail to capture the real execution statistics of production cloud workloads. Although some cloud database vendors have recently released real workload traces, these traces alone do not qualify as benchmarks, as they typically lack essential components like the original SQL queries and their underlying databases. To overcome this limitation, this paper introduces a new problem of workload synthesis with real statistics, which aims to generate synthetic workloads that closely approximate real execution statistics, including key performance metrics and operator distributions, in real cloud workloads. To address this problem, we propose PBench, a novel workload synthesizer that constructs synthetic workloads by judiciously selecting and combining workload components (i.e., queries and databases) from existing benchmarks. This paper studies the key challenges in PBench. First, we address the challenge of balancing performance metrics and operator distributions by introducing a multi-objective optimization-based component selection method. Second, to capture the temporal dynamics of real workloads, we design a timestamp assignment method that progressively refines workload timestamps. Third, to handle the disparity between the original workload and the candidate workload, we propose a component augmentation approach that leverages large language models (LLMs) to generate additional workload components while maintaining statistical fidelity. We evaluate PBench on real cloud workload traces, demonstrating that it reduces approximation error by up to 6x compared to state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16379
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PBench: Workload Synthesizer with Real Statistics for Cloud Analytics Benchmarking
Zhou, Yan
Liu, Chunwei
Urgaonkar, Bhuvan
Wang, Zhengle
Mueller, Magnus
Zhang, Chao
Zhang, Songyue
Pfeil, Pascal
Horn, Dominik
Liu, Zhengchun
Pagano, Davide
Kraska, Tim
Madden, Samuel
Fan, Ju
Databases
Cloud service providers commonly use standard benchmarks like TPC-H and TPC-DS to evaluate and optimize cloud data analytics systems. However, these benchmarks rely on fixed query patterns and fail to capture the real execution statistics of production cloud workloads. Although some cloud database vendors have recently released real workload traces, these traces alone do not qualify as benchmarks, as they typically lack essential components like the original SQL queries and their underlying databases. To overcome this limitation, this paper introduces a new problem of workload synthesis with real statistics, which aims to generate synthetic workloads that closely approximate real execution statistics, including key performance metrics and operator distributions, in real cloud workloads. To address this problem, we propose PBench, a novel workload synthesizer that constructs synthetic workloads by judiciously selecting and combining workload components (i.e., queries and databases) from existing benchmarks. This paper studies the key challenges in PBench. First, we address the challenge of balancing performance metrics and operator distributions by introducing a multi-objective optimization-based component selection method. Second, to capture the temporal dynamics of real workloads, we design a timestamp assignment method that progressively refines workload timestamps. Third, to handle the disparity between the original workload and the candidate workload, we propose a component augmentation approach that leverages large language models (LLMs) to generate additional workload components while maintaining statistical fidelity. We evaluate PBench on real cloud workload traces, demonstrating that it reduces approximation error by up to 6x compared to state-of-the-art methods.
title PBench: Workload Synthesizer with Real Statistics for Cloud Analytics Benchmarking
topic Databases
url https://arxiv.org/abs/2506.16379