RASL: Retrieval Augmented Schema Linking for Massive Database Text-to-SQL

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Eben, Jeffrey, Ahmad, Aitzaz, Lau, Stephen
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911084813746176
author Eben, Jeffrey
Ahmad, Aitzaz
Lau, Stephen
author_facet Eben, Jeffrey
Ahmad, Aitzaz
Lau, Stephen
contents Despite advances in large language model (LLM)-based natural language interfaces for databases, scaling to enterprise-level data catalogs remains an under-explored challenge. Prior works addressing this challenge rely on domain-specific fine-tuning - complicating deployment - and fail to leverage important semantic context contained within database metadata. To address these limitations, we introduce a component-based retrieval architecture that decomposes database schemas and metadata into discrete semantic units, each separately indexed for targeted retrieval. Our approach prioritizes effective table identification while leveraging column-level information, ensuring the total number of retrieved tables remains within a manageable context budget. Experiments demonstrate that our method maintains high recall and accuracy, with our system outperforming baselines over massive databases with varying structure and available metadata. Our solution enables practical text-to-SQL systems deployable across diverse enterprise settings without specialized fine-tuning, addressing a critical scalability gap in natural language database interfaces.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23104
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RASL: Retrieval Augmented Schema Linking for Massive Database Text-to-SQL
Eben, Jeffrey
Ahmad, Aitzaz
Lau, Stephen
Computation and Language
Artificial Intelligence
Machine Learning
Despite advances in large language model (LLM)-based natural language interfaces for databases, scaling to enterprise-level data catalogs remains an under-explored challenge. Prior works addressing this challenge rely on domain-specific fine-tuning - complicating deployment - and fail to leverage important semantic context contained within database metadata. To address these limitations, we introduce a component-based retrieval architecture that decomposes database schemas and metadata into discrete semantic units, each separately indexed for targeted retrieval. Our approach prioritizes effective table identification while leveraging column-level information, ensuring the total number of retrieved tables remains within a manageable context budget. Experiments demonstrate that our method maintains high recall and accuracy, with our system outperforming baselines over massive databases with varying structure and available metadata. Our solution enables practical text-to-SQL systems deployable across diverse enterprise settings without specialized fine-tuning, addressing a critical scalability gap in natural language database interfaces.
title RASL: Retrieval Augmented Schema Linking for Massive Database Text-to-SQL
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2507.23104