CSR-RAG: An Efficient Retrieval System for Text-to-SQL on the Enterprise Scale

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singh, Rajpreet, Boškov, Novak, Drabeck, Lawrence, Gudal, Aditya, Khan, Manzoor A.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914245364416512
author Singh, Rajpreet
Boškov, Novak
Drabeck, Lawrence
Gudal, Aditya
Khan, Manzoor A.
author_facet Singh, Rajpreet
Boškov, Novak
Drabeck, Lawrence
Gudal, Aditya
Khan, Manzoor A.
contents Natural language to SQL translation (Text-to-SQL) is one of the long-standing problems that has recently benefited from advances in Large Language Models (LLMs). While most academic Text-to-SQL benchmarks request schema description as a part of natural language input, enterprise-scale applications often require table retrieval before SQL query generation. To address this need, we propose a novel hybrid Retrieval Augmented Generation (RAG) system consisting of contextual, structural, and relational retrieval (CSR-RAG) to achieve computationally efficient yet sufficiently accurate retrieval for enterprise-scale databases. Through extensive enterprise benchmarks, we demonstrate that CSR-RAG achieves up to 40% precision and over 80% recall while incurring a negligible average query generation latency of only 30ms on commodity data center hardware, which makes it appropriate for modern LLM-based enterprise-scale systems.
format Preprint
id arxiv_https___arxiv_org_abs_2601_06564
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CSR-RAG: An Efficient Retrieval System for Text-to-SQL on the Enterprise Scale
Singh, Rajpreet
Boškov, Novak
Drabeck, Lawrence
Gudal, Aditya
Khan, Manzoor A.
Computation and Language
Natural language to SQL translation (Text-to-SQL) is one of the long-standing problems that has recently benefited from advances in Large Language Models (LLMs). While most academic Text-to-SQL benchmarks request schema description as a part of natural language input, enterprise-scale applications often require table retrieval before SQL query generation. To address this need, we propose a novel hybrid Retrieval Augmented Generation (RAG) system consisting of contextual, structural, and relational retrieval (CSR-RAG) to achieve computationally efficient yet sufficiently accurate retrieval for enterprise-scale databases. Through extensive enterprise benchmarks, we demonstrate that CSR-RAG achieves up to 40% precision and over 80% recall while incurring a negligible average query generation latency of only 30ms on commodity data center hardware, which makes it appropriate for modern LLM-based enterprise-scale systems.
title CSR-RAG: An Efficient Retrieval System for Text-to-SQL on the Enterprise Scale
topic Computation and Language
url https://arxiv.org/abs/2601.06564