Lightweight Transformers for Zero-Shot and Fine-Tuned Text-to-SQL Generation Using Spider

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Seth, Chirag, Singh, Utkarsh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908480316637184
author Seth, Chirag
Singh, Utkarsh
author_facet Seth, Chirag
Singh, Utkarsh
contents Text-to-SQL translation enables non-expert users to query relational databases using natural language, with applications in education and business intelligence. This study evaluates three lightweight transformer models - T5-Small, BART-Small, and GPT-2 - on the Spider dataset, focusing on low-resource settings. We developed a reusable, model-agnostic pipeline that tailors schema formatting to each model's architecture, training them across 1000 to 5000 iterations and evaluating on 1000 test samples using Logical Form Accuracy (LFAcc), BLEU, and Exact Match (EM) metrics. Fine-tuned T5-Small achieves the highest LFAcc (27.8%), outperforming BART-Small (23.98%) and GPT-2 (20.1%), highlighting encoder-decoder models' superiority in schema-aware SQL generation. Despite resource constraints limiting performance, our pipeline's modularity supports future enhancements, such as advanced schema linking or alternative base models. This work underscores the potential of compact transformers for accessible text-to-SQL solutions in resource-scarce environments.
format Preprint
id arxiv_https___arxiv_org_abs_2508_04623
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Lightweight Transformers for Zero-Shot and Fine-Tuned Text-to-SQL Generation Using Spider
Seth, Chirag
Singh, Utkarsh
Computation and Language
Information Retrieval
68T50 % Natural language processing (in Computer Science)
I.2.7; H.2.3
Text-to-SQL translation enables non-expert users to query relational databases using natural language, with applications in education and business intelligence. This study evaluates three lightweight transformer models - T5-Small, BART-Small, and GPT-2 - on the Spider dataset, focusing on low-resource settings. We developed a reusable, model-agnostic pipeline that tailors schema formatting to each model's architecture, training them across 1000 to 5000 iterations and evaluating on 1000 test samples using Logical Form Accuracy (LFAcc), BLEU, and Exact Match (EM) metrics. Fine-tuned T5-Small achieves the highest LFAcc (27.8%), outperforming BART-Small (23.98%) and GPT-2 (20.1%), highlighting encoder-decoder models' superiority in schema-aware SQL generation. Despite resource constraints limiting performance, our pipeline's modularity supports future enhancements, such as advanced schema linking or alternative base models. This work underscores the potential of compact transformers for accessible text-to-SQL solutions in resource-scarce environments.
title Lightweight Transformers for Zero-Shot and Fine-Tuned Text-to-SQL Generation Using Spider
topic Computation and Language
Information Retrieval
68T50 % Natural language processing (in Computer Science)
I.2.7; H.2.3
url https://arxiv.org/abs/2508.04623