Evaluating Learned Query Performance Prediction Models at LinkedIn: Challenges, Opportunities, and Findings

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Song, Chujun, Bouguerra, Slim, Krogen, Erik, Abadi, Daniel
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913806586740736
author Song, Chujun
Bouguerra, Slim
Krogen, Erik
Abadi, Daniel
author_facet Song, Chujun
Bouguerra, Slim
Krogen, Erik
Abadi, Daniel
contents Recent advancements in learning-based query performance prediction models have demonstrated remarkable efficacy. However, these models are predominantly validated using synthetic datasets focused on cardinality or latency estimations. This paper explores the application of these models to LinkedIn's complex real-world OLAP queries executed on Trino, addressing four primary research questions: (1) How do these models perform on real-world industrial data with limited information? (2) Can these models generalize to new tasks, such as CPU time prediction and classification? (3) What additional information available from the query plan could be utilized by these models to enhance their performance? (4) What are the theoretical performance limits of these models given the available data? To address these questions, we evaluate several models-including TLSTM, TCNN, QueryFormer, and XGBoost, against the industrial query workload at LinkedIn, and extend our analysis to CPU time regression and classification tasks. We also propose a multi-task learning approach to incorporate underutilized operator-level metrics that could enhance model understanding. Additionally, we empirically analyze the inherent upper bound that can be achieved from the models.
format Preprint
id arxiv_https___arxiv_org_abs_2504_17181
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Learned Query Performance Prediction Models at LinkedIn: Challenges, Opportunities, and Findings
Song, Chujun
Bouguerra, Slim
Krogen, Erik
Abadi, Daniel
Databases
H.2.4
Recent advancements in learning-based query performance prediction models have demonstrated remarkable efficacy. However, these models are predominantly validated using synthetic datasets focused on cardinality or latency estimations. This paper explores the application of these models to LinkedIn's complex real-world OLAP queries executed on Trino, addressing four primary research questions: (1) How do these models perform on real-world industrial data with limited information? (2) Can these models generalize to new tasks, such as CPU time prediction and classification? (3) What additional information available from the query plan could be utilized by these models to enhance their performance? (4) What are the theoretical performance limits of these models given the available data? To address these questions, we evaluate several models-including TLSTM, TCNN, QueryFormer, and XGBoost, against the industrial query workload at LinkedIn, and extend our analysis to CPU time regression and classification tasks. We also propose a multi-task learning approach to incorporate underutilized operator-level metrics that could enhance model understanding. Additionally, we empirically analyze the inherent upper bound that can be achieved from the models.
title Evaluating Learned Query Performance Prediction Models at LinkedIn: Challenges, Opportunities, and Findings
topic Databases
H.2.4
url https://arxiv.org/abs/2504.17181