DistJoin: A Decoupled Join Cardinality Estimator based on Adaptive Neural Predicate Modulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Kaixin, Wang, Hongzhi, Li, Ziqi, Lu, Yabin, Li, Yingze, Yan, Yu, Guan, Yiming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911143479476224
author Zhang, Kaixin
Wang, Hongzhi
Li, Ziqi
Lu, Yabin
Li, Yingze
Yan, Yu
Guan, Yiming
author_facet Zhang, Kaixin
Wang, Hongzhi
Li, Ziqi
Lu, Yabin
Li, Yingze
Yan, Yu
Guan, Yiming
contents Research on learned cardinality estimation has made significant progress in recent years. However, existing methods still face distinct challenges that hinder their practical deployment in production environments. We define these challenges as the ``Trilemma of Cardinality Estimation'', where learned cardinality estimation methods struggle to balance generality, accuracy, and updatability. To address these challenges, we introduce DistJoin, a join cardinality estimator based on efficient distribution prediction using multi-autoregressive models. Our contributions are threefold: (1) We propose a method to estimate join cardinality by leveraging the probability distributions of individual tables in a decoupled manner. (2) To meet the requirements of efficiency for DistJoin, we develop Adaptive Neural Predicate Modulation (ANPM), a high-throughput distribution estimation model. (3) We demonstrate that an existing similar approach suffers from variance accumulation issues by formal variance analysis. To mitigate this problem, DistJoin employs a selectivity-based approach to infer join cardinality, effectively reducing variance. In summary, DistJoin not only represents the first data-driven method to support both equi and non-equi joins simultaneously but also demonstrates superior accuracy while enabling fast and flexible updates. The experimental results demonstrate that DistJoin achieves the highest accuracy, robustness to data updates, generality, and comparable update and inference speed relative to existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2503_08994
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DistJoin: A Decoupled Join Cardinality Estimator based on Adaptive Neural Predicate Modulation
Zhang, Kaixin
Wang, Hongzhi
Li, Ziqi
Lu, Yabin
Li, Yingze
Yan, Yu
Guan, Yiming
Databases
Artificial Intelligence
Research on learned cardinality estimation has made significant progress in recent years. However, existing methods still face distinct challenges that hinder their practical deployment in production environments. We define these challenges as the ``Trilemma of Cardinality Estimation'', where learned cardinality estimation methods struggle to balance generality, accuracy, and updatability. To address these challenges, we introduce DistJoin, a join cardinality estimator based on efficient distribution prediction using multi-autoregressive models. Our contributions are threefold: (1) We propose a method to estimate join cardinality by leveraging the probability distributions of individual tables in a decoupled manner. (2) To meet the requirements of efficiency for DistJoin, we develop Adaptive Neural Predicate Modulation (ANPM), a high-throughput distribution estimation model. (3) We demonstrate that an existing similar approach suffers from variance accumulation issues by formal variance analysis. To mitigate this problem, DistJoin employs a selectivity-based approach to infer join cardinality, effectively reducing variance. In summary, DistJoin not only represents the first data-driven method to support both equi and non-equi joins simultaneously but also demonstrates superior accuracy while enabling fast and flexible updates. The experimental results demonstrate that DistJoin achieves the highest accuracy, robustness to data updates, generality, and comparable update and inference speed relative to existing methods.
title DistJoin: A Decoupled Join Cardinality Estimator based on Adaptive Neural Predicate Modulation
topic Databases
Artificial Intelligence
url https://arxiv.org/abs/2503.08994