Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Yue, Mimno, David, Jo, Unso Eun Seo
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2603.25568
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912983868768256
author Li, Yue
Mimno, David
Jo, Unso Eun Seo
author_facet Li, Yue
Mimno, David
Jo, Unso Eun Seo
contents Translating natural language to SQL for data retrieval has become more accessible thanks to code generation LLMs. But how hard is it to generate SQL code? While databases can become unbounded in complexity, the complexity of queries is bounded by real life utility and human needs. With a sample of 376 databases, we show that SQL queries, as translations of natural language questions are finite in practical complexity. There is no clear monotonic relationship between increases in database table count and increases in complexity of SQL queries. In their template forms, SQL queries follow a Power Law-like distribution of frequency where 70% of our tested queries can be covered with just 13% of all template types, indicating that the high majority of SQL queries are predictable. This suggests that while LLMs for code generation can be useful, in the domain of database access, they may be operating in a narrow, highly formulaic space where templates could be safer, cheaper, and auditable.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25568
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Are LLMs Overkill for Databases?: A Study on the Finiteness of SQL
Li, Yue
Mimno, David
Jo, Unso Eun Seo
Databases
Artificial Intelligence
Translating natural language to SQL for data retrieval has become more accessible thanks to code generation LLMs. But how hard is it to generate SQL code? While databases can become unbounded in complexity, the complexity of queries is bounded by real life utility and human needs. With a sample of 376 databases, we show that SQL queries, as translations of natural language questions are finite in practical complexity. There is no clear monotonic relationship between increases in database table count and increases in complexity of SQL queries. In their template forms, SQL queries follow a Power Law-like distribution of frequency where 70% of our tested queries can be covered with just 13% of all template types, indicating that the high majority of SQL queries are predictable. This suggests that while LLMs for code generation can be useful, in the domain of database access, they may be operating in a narrow, highly formulaic space where templates could be safer, cheaper, and auditable.
title Are LLMs Overkill for Databases?: A Study on the Finiteness of SQL
topic Databases
Artificial Intelligence
url https://arxiv.org/abs/2603.25568