SpotIt: Evaluating Text-to-SQL Evaluation with Formal Verification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Klopfenstein, Rocky, He, Yang, Tremante, Andrew, Wang, Yuepeng, Narodytska, Nina, Wu, Haoze
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912942414364672
author Klopfenstein, Rocky
He, Yang
Tremante, Andrew
Wang, Yuepeng
Narodytska, Nina
Wu, Haoze
author_facet Klopfenstein, Rocky
He, Yang
Tremante, Andrew
Wang, Yuepeng
Narodytska, Nina
Wu, Haoze
contents Community-driven Text-to-SQL evaluation platforms play a pivotal role in tracking the state of the art of Text-to-SQL performance. The reliability of the evaluation process is critical for driving progress in the field. Current evaluation methods are largely test-based, which involves comparing the execution results of a generated SQL query and a human-labeled ground-truth on a static test database. Such an evaluation is optimistic, as two queries can coincidentally produce the same output on the test database while actually being different. In this work, we propose a new alternative evaluation pipeline, called SpotIt, where a formal bounded equivalence verification engine actively searches for a database that differentiates the generated and ground-truth SQL queries. We develop techniques to extend existing verifiers to support a richer SQL subset relevant to Text-to-SQL. A performance evaluation of ten Text-to-SQL methods on the high-profile BIRD dataset suggests that test-based methods can often overlook differences between the generated query and the ground-truth. Further analysis of the verification results reveals a more complex picture of the current Text-to-SQL evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26840
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpotIt: Evaluating Text-to-SQL Evaluation with Formal Verification
Klopfenstein, Rocky
He, Yang
Tremante, Andrew
Wang, Yuepeng
Narodytska, Nina
Wu, Haoze
Databases
Artificial Intelligence
Formal Languages and Automata Theory
Logic in Computer Science
Community-driven Text-to-SQL evaluation platforms play a pivotal role in tracking the state of the art of Text-to-SQL performance. The reliability of the evaluation process is critical for driving progress in the field. Current evaluation methods are largely test-based, which involves comparing the execution results of a generated SQL query and a human-labeled ground-truth on a static test database. Such an evaluation is optimistic, as two queries can coincidentally produce the same output on the test database while actually being different. In this work, we propose a new alternative evaluation pipeline, called SpotIt, where a formal bounded equivalence verification engine actively searches for a database that differentiates the generated and ground-truth SQL queries. We develop techniques to extend existing verifiers to support a richer SQL subset relevant to Text-to-SQL. A performance evaluation of ten Text-to-SQL methods on the high-profile BIRD dataset suggests that test-based methods can often overlook differences between the generated query and the ground-truth. Further analysis of the verification results reveals a more complex picture of the current Text-to-SQL evaluation.
title SpotIt: Evaluating Text-to-SQL Evaluation with Formal Verification
topic Databases
Artificial Intelligence
Formal Languages and Automata Theory
Logic in Computer Science
url https://arxiv.org/abs/2510.26840