Knowledge Base Construction for Knowledge-Augmented Text-to-SQL

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Baek, Jinheon, Samulowitz, Horst, Hassanzadeh, Oktie, Subramanian, Dharmashankar, Shirai, Sola, Gliozzo, Alfio, Bhattacharjya, Debarun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916764013559808
author Baek, Jinheon
Samulowitz, Horst
Hassanzadeh, Oktie
Subramanian, Dharmashankar
Shirai, Sola
Gliozzo, Alfio
Bhattacharjya, Debarun
author_facet Baek, Jinheon
Samulowitz, Horst
Hassanzadeh, Oktie
Subramanian, Dharmashankar
Shirai, Sola
Gliozzo, Alfio
Bhattacharjya, Debarun
contents Text-to-SQL aims to translate natural language queries into SQL statements, which is practical as it enables anyone to easily retrieve the desired information from databases. Recently, many existing approaches tackle this problem with Large Language Models (LLMs), leveraging their strong capability in understanding user queries and generating corresponding SQL code. Yet, the parametric knowledge in LLMs might be limited to covering all the diverse and domain-specific queries that require grounding in various database schemas, which makes generated SQLs less accurate oftentimes. To tackle this, we propose constructing the knowledge base for text-to-SQL, a foundational source of knowledge, from which we retrieve and generate the necessary knowledge for given queries. In particular, unlike existing approaches that either manually annotate knowledge or generate only a few pieces of knowledge for each query, our knowledge base is comprehensive, which is constructed based on a combination of all the available questions and their associated database schemas along with their relevant knowledge, and can be reused for unseen databases from different datasets and domains. We validate our approach on multiple text-to-SQL datasets, considering both the overlapping and non-overlapping database scenarios, where it outperforms relevant baselines substantially.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22096
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Knowledge Base Construction for Knowledge-Augmented Text-to-SQL
Baek, Jinheon
Samulowitz, Horst
Hassanzadeh, Oktie
Subramanian, Dharmashankar
Shirai, Sola
Gliozzo, Alfio
Bhattacharjya, Debarun
Computation and Language
Artificial Intelligence
Machine Learning
Text-to-SQL aims to translate natural language queries into SQL statements, which is practical as it enables anyone to easily retrieve the desired information from databases. Recently, many existing approaches tackle this problem with Large Language Models (LLMs), leveraging their strong capability in understanding user queries and generating corresponding SQL code. Yet, the parametric knowledge in LLMs might be limited to covering all the diverse and domain-specific queries that require grounding in various database schemas, which makes generated SQLs less accurate oftentimes. To tackle this, we propose constructing the knowledge base for text-to-SQL, a foundational source of knowledge, from which we retrieve and generate the necessary knowledge for given queries. In particular, unlike existing approaches that either manually annotate knowledge or generate only a few pieces of knowledge for each query, our knowledge base is comprehensive, which is constructed based on a combination of all the available questions and their associated database schemas along with their relevant knowledge, and can be reused for unseen databases from different datasets and domains. We validate our approach on multiple text-to-SQL datasets, considering both the overlapping and non-overlapping database scenarios, where it outperforms relevant baselines substantially.
title Knowledge Base Construction for Knowledge-Augmented Text-to-SQL
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.22096