Can Large Language Models be a Cardinality Estimator? An Empirical study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Liangzu, Wang, Yiyan, Wu, Yinjun, Su, Runze, Chang, Zhuo, Wu, Peizhi, Chen, Jianjun, Jiang, Fuxin, Shi, Rui, Cui, Bin, Zhang, Tieying
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914431520210944
author Liu, Liangzu
Wang, Yiyan
Wu, Yinjun
Su, Runze
Chang, Zhuo
Wu, Peizhi
Chen, Jianjun
Jiang, Fuxin
Shi, Rui
Cui, Bin
Zhang, Tieying
author_facet Liu, Liangzu
Wang, Yiyan
Wu, Yinjun
Su, Runze
Chang, Zhuo
Wu, Peizhi
Chen, Jianjun
Jiang, Fuxin
Shi, Rui
Cui, Bin
Zhang, Tieying
contents Cardinality estimation (CardEst) still remains a challenging problem for DBMS. Recent years have witnessed the success of ML-based cardinality estimators in outperforming traditional methods. However, these solutions suffer from poor generalizability to new data or query distribution, inability to handle complex queries, and substantial data preparation overhead, thus preventing their wide adoption in the real-world DBMS. Some recent efforts have been dedicated to addressing some but not all of these issues. We notice that the recent emerging Large Language Models (LLMs) have shown their remarkable generalizability to unseen tasks, capabilities to understand complex programs, and power to perform data-efficient fine-tuning. In light of this, we propose to leverage LLMs to mitigate the above issues. Specifically, we carefully craft prompts, and subsequently perform fine-tuning and self-correction during inference with LLMs for CardEst task. We then extensively evaluate LLMs' in-distribution and out-of-distribution generalizability, feasibility to support complex queries, and training data efficiency during fine-tuning LLMs on pre-training datasets. The results suggest that LLMs outperform the state-of-the-art in almost all settings, thus indicating their potential for the CardEst task. We further measure the end-to-end query execution time in DBMS by using the estimated cardinalities of LLMs in some practical settings, which suggests that the inference overhead of LLMs can be outweighed by the benefits brought by LLMs for CardEst.
format Preprint
id arxiv_https___arxiv_org_abs_2603_28080
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Can Large Language Models be a Cardinality Estimator? An Empirical study
Liu, Liangzu
Wang, Yiyan
Wu, Yinjun
Su, Runze
Chang, Zhuo
Wu, Peizhi
Chen, Jianjun
Jiang, Fuxin
Shi, Rui
Cui, Bin
Zhang, Tieying
Databases
Cardinality estimation (CardEst) still remains a challenging problem for DBMS. Recent years have witnessed the success of ML-based cardinality estimators in outperforming traditional methods. However, these solutions suffer from poor generalizability to new data or query distribution, inability to handle complex queries, and substantial data preparation overhead, thus preventing their wide adoption in the real-world DBMS. Some recent efforts have been dedicated to addressing some but not all of these issues. We notice that the recent emerging Large Language Models (LLMs) have shown their remarkable generalizability to unseen tasks, capabilities to understand complex programs, and power to perform data-efficient fine-tuning. In light of this, we propose to leverage LLMs to mitigate the above issues. Specifically, we carefully craft prompts, and subsequently perform fine-tuning and self-correction during inference with LLMs for CardEst task. We then extensively evaluate LLMs' in-distribution and out-of-distribution generalizability, feasibility to support complex queries, and training data efficiency during fine-tuning LLMs on pre-training datasets. The results suggest that LLMs outperform the state-of-the-art in almost all settings, thus indicating their potential for the CardEst task. We further measure the end-to-end query execution time in DBMS by using the estimated cardinalities of LLMs in some practical settings, which suggests that the inference overhead of LLMs can be outweighed by the benefits brought by LLMs for CardEst.
title Can Large Language Models be a Cardinality Estimator? An Empirical study
topic Databases
url https://arxiv.org/abs/2603.28080