Text-to-SQL based on Large Language Models and Database Keyword Search

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Nascimento, Eduardo R., Avila, Caio Viktor S., Izquierdo, Yenier T., García, Grettel M., Andrade, Lucas Feijó L., Facina, Michelle S. P., Lemos, Melissa, Casanova, Marco A.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917900688818176
author Nascimento, Eduardo R.
Avila, Caio Viktor S.
Izquierdo, Yenier T.
García, Grettel M.
Andrade, Lucas Feijó L.
Facina, Michelle S. P.
Lemos, Melissa
Casanova, Marco A.
author_facet Nascimento, Eduardo R.
Avila, Caio Viktor S.
Izquierdo, Yenier T.
García, Grettel M.
Andrade, Lucas Feijó L.
Facina, Michelle S. P.
Lemos, Melissa
Casanova, Marco A.
contents Text-to-SQL prompt strategies based on Large Language Models (LLMs) achieve remarkable performance on well-known benchmarks. However, when applied to real-world databases, their performance is significantly less than for these benchmarks, especially for Natural Language (NL) questions requiring complex filters and joins to be processed. This paper then proposes a strategy to compile NL questions into SQL queries that incorporates a dynamic few-shot examples strategy and leverages the services provided by a database keyword search (KwS) platform. The paper details how the precision and recall of the schema-linking process are improved with the help of the examples provided and the keyword-matching service that the KwS platform offers. Then, it shows how the KwS platform can be used to synthesize a view that captures the joins required to process an input NL question and thereby simplify the SQL query compilation step. The paper includes experiments with a real-world relational database to assess the performance of the proposed strategy. The experiments suggest that the strategy achieves an accuracy on the real-world relational database that surpasses state-of-the-art approaches. The paper concludes by discussing the results obtained.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13594
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Text-to-SQL based on Large Language Models and Database Keyword Search
Nascimento, Eduardo R.
Avila, Caio Viktor S.
Izquierdo, Yenier T.
García, Grettel M.
Andrade, Lucas Feijó L.
Facina, Michelle S. P.
Lemos, Melissa
Casanova, Marco A.
Databases
Artificial Intelligence
68T50
H.2.3; I.2.7
Text-to-SQL prompt strategies based on Large Language Models (LLMs) achieve remarkable performance on well-known benchmarks. However, when applied to real-world databases, their performance is significantly less than for these benchmarks, especially for Natural Language (NL) questions requiring complex filters and joins to be processed. This paper then proposes a strategy to compile NL questions into SQL queries that incorporates a dynamic few-shot examples strategy and leverages the services provided by a database keyword search (KwS) platform. The paper details how the precision and recall of the schema-linking process are improved with the help of the examples provided and the keyword-matching service that the KwS platform offers. Then, it shows how the KwS platform can be used to synthesize a view that captures the joins required to process an input NL question and thereby simplify the SQL query compilation step. The paper includes experiments with a real-world relational database to assess the performance of the proposed strategy. The experiments suggest that the strategy achieves an accuracy on the real-world relational database that surpasses state-of-the-art approaches. The paper concludes by discussing the results obtained.
title Text-to-SQL based on Large Language Models and Database Keyword Search
topic Databases
Artificial Intelligence
68T50
H.2.3; I.2.7
url https://arxiv.org/abs/2501.13594