SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Safarzadeh, Mohammadtaher, Patel, Hitesh Laxmichand, Orojlooyjadid, Afshin, Horwood, Graham, Roth, Dan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915944881717248
author Safarzadeh, Mohammadtaher
Patel, Hitesh Laxmichand
Orojlooyjadid, Afshin
Horwood, Graham
Roth, Dan
author_facet Safarzadeh, Mohammadtaher
Patel, Hitesh Laxmichand
Orojlooyjadid, Afshin
Horwood, Graham
Roth, Dan
contents Large language models (LLMs) have achieved strong performance on natural language to SQL (NL2SQL) benchmarks, yet their reported accuracy may be inflated by contamination from benchmark queries or structurally similar patterns seen during training. We introduce SPENCE (Syntactic Probing and Evaluation of NL2SQL Contamination Effects), a controlled syntactic probing framework for detecting and quantifying such contamination. SPENCE systematically generates syntactic variants of test queries for four widely used NL2SQL datasets-Spider, SParC, CoSQL, and the newer BIRD benchmark. We use SPENCE to evaluate multiple high-capacity LLMs under execution-based scoring. For each model, we measure changes in execution accuracy across increasing levels of syntactic divergence and quantify rank sensitivity using Kendall's tau with bootstrap confidence intervals. By aligning these robustness trends with benchmark release dates, we observe a clear temporal gradient: older benchmarks such as Spider exhibit the strongest negative values and thus the highest likelihood of training leakage, whereas the more recent BIRD dataset shows minimal sensitivity and appears largely uncontaminated. Together, these findings highlight the importance of temporally contextualized, syntactic-probing evaluation for trustworthy NL2SQL benchmarking.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17771
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks
Safarzadeh, Mohammadtaher
Patel, Hitesh Laxmichand
Orojlooyjadid, Afshin
Horwood, Graham
Roth, Dan
Computation and Language
Artificial Intelligence
Databases
Large language models (LLMs) have achieved strong performance on natural language to SQL (NL2SQL) benchmarks, yet their reported accuracy may be inflated by contamination from benchmark queries or structurally similar patterns seen during training. We introduce SPENCE (Syntactic Probing and Evaluation of NL2SQL Contamination Effects), a controlled syntactic probing framework for detecting and quantifying such contamination. SPENCE systematically generates syntactic variants of test queries for four widely used NL2SQL datasets-Spider, SParC, CoSQL, and the newer BIRD benchmark. We use SPENCE to evaluate multiple high-capacity LLMs under execution-based scoring. For each model, we measure changes in execution accuracy across increasing levels of syntactic divergence and quantify rank sensitivity using Kendall's tau with bootstrap confidence intervals. By aligning these robustness trends with benchmark release dates, we observe a clear temporal gradient: older benchmarks such as Spider exhibit the strongest negative values and thus the highest likelihood of training leakage, whereas the more recent BIRD dataset shows minimal sensitivity and appears largely uncontaminated. Together, these findings highlight the importance of temporally contextualized, syntactic-probing evaluation for trustworthy NL2SQL benchmarking.
title SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks
topic Computation and Language
Artificial Intelligence
Databases
url https://arxiv.org/abs/2604.17771