SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhou, Yixi, Zhang, Fan, Guo, Zhiqiao, Chen, Yu, Zhang, Haipeng, Nakov, Preslav, Xie, Zhuohan
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914457457786880
author Zhou, Yixi
Zhang, Fan
Guo, Zhiqiao
Chen, Yu
Zhang, Haipeng
Nakov, Preslav
Xie, Zhuohan
author_facet Zhou, Yixi
Zhang, Fan
Guo, Zhiqiao
Chen, Yu
Zhang, Haipeng
Nakov, Preslav
Xie, Zhuohan
contents Despite strong performance on Text-to-SQL benchmarks, it remains unclear whether LLM-generated SQL programs are structurally reliable. In this work, we investigate the structural behavior of LLM-generated SQL queries and introduce SQLStructEval, a framework for analyzing program structures through canonical abstract syntax tree (AST) representations. Our experiments on the Spider benchmark show that modern LLMs often produce structurally diverse queries for the same input, even when execution results are correct, and that such variance is frequently triggered by surface-level input changes such as paraphrases or schema presentation. We further show that generating queries in a structured space via a compile-style pipeline can improve both execution accuracy and structural consistency. These findings suggest that structural reliability is a critical yet overlooked dimension for evaluating LLM-based program generation systems. Our code is available at https://anonymous.4open.science/r/StructEval-2435.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06736
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation
Zhou, Yixi
Zhang, Fan
Guo, Zhiqiao
Chen, Yu
Zhang, Haipeng
Nakov, Preslav
Xie, Zhuohan
Computation and Language
Databases
Despite strong performance on Text-to-SQL benchmarks, it remains unclear whether LLM-generated SQL programs are structurally reliable. In this work, we investigate the structural behavior of LLM-generated SQL queries and introduce SQLStructEval, a framework for analyzing program structures through canonical abstract syntax tree (AST) representations. Our experiments on the Spider benchmark show that modern LLMs often produce structurally diverse queries for the same input, even when execution results are correct, and that such variance is frequently triggered by surface-level input changes such as paraphrases or schema presentation. We further show that generating queries in a structured space via a compile-style pipeline can improve both execution accuracy and structural consistency. These findings suggest that structural reliability is a critical yet overlooked dimension for evaluating LLM-based program generation systems. Our code is available at https://anonymous.4open.science/r/StructEval-2435.
title SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation
topic Computation and Language
Databases
url https://arxiv.org/abs/2604.06736