Staff View: :: Library Catalog

Saved in:

Bibliographic Details
Main Authors:	Kim, Heegyu, Jeon, Taeyang, Choi, Seunghwan, Choi, Seungtaek, Cho, Hyunsouk
Format:	Preprint
Published:	2024
Subjects:	Computation and Language Information Retrieval Machine Learning
Online Access:	https://arxiv.org/abs/2409.19014
Tags:	Add Tag No Tags, Be the first to tag this record!

_version_	1866909367410884608
author	Kim, Heegyu Jeon, Taeyang Choi, Seunghwan Choi, Seungtaek Cho, Hyunsouk
author_facet	Kim, Heegyu Jeon, Taeyang Choi, Seunghwan Choi, Seungtaek Cho, Hyunsouk
contents	Text-to-SQL systems have become crucial for translating natural language into SQL queries in various industries, enabling non-technical users to perform complex data operations. The need for accurate evaluation methods has increased as these systems have grown more sophisticated. However, the Execution Accuracy (EX), the most prevalent evaluation metric, still shows many false positives and negatives. Thus, this paper introduces FLEX (False-Less EXecution), a novel approach to evaluating text-to-SQL systems using large language models (LLMs) to emulate human expert-level evaluation of SQL queries. Our metric improves agreement with human experts (from 62 to 87.04 in Cohen's kappa) with comprehensive context and sophisticated criteria. Our extensive experiments yield several key insights: (1) Models' performance increases by over 2.6 points on average, substantially affecting rankings on Spider and BIRD benchmarks; (2) The underestimation of models in EX primarily stems from annotation quality issues; and (3) Model performance on particularly challenging questions tends to be overestimated. This work contributes to a more accurate and nuanced evaluation of text-to-SQL systems, potentially reshaping our understanding of state-of-the-art performance in this field.
format	Preprint
id	arxiv_https___arxiv_org_abs_2409_19014
institution	arXiv
publishDate	2024
record_format	arxiv
spellingShingle	FLEX: Expert-level False-Less EXecution Metric for Reliable Text-to-SQL Benchmark Kim, Heegyu Jeon, Taeyang Choi, Seunghwan Choi, Seungtaek Cho, Hyunsouk Computation and Language Information Retrieval Machine Learning Text-to-SQL systems have become crucial for translating natural language into SQL queries in various industries, enabling non-technical users to perform complex data operations. The need for accurate evaluation methods has increased as these systems have grown more sophisticated. However, the Execution Accuracy (EX), the most prevalent evaluation metric, still shows many false positives and negatives. Thus, this paper introduces FLEX (False-Less EXecution), a novel approach to evaluating text-to-SQL systems using large language models (LLMs) to emulate human expert-level evaluation of SQL queries. Our metric improves agreement with human experts (from 62 to 87.04 in Cohen's kappa) with comprehensive context and sophisticated criteria. Our extensive experiments yield several key insights: (1) Models' performance increases by over 2.6 points on average, substantially affecting rankings on Spider and BIRD benchmarks; (2) The underestimation of models in EX primarily stems from annotation quality issues; and (3) Model performance on particularly challenging questions tends to be overestimated. This work contributes to a more accurate and nuanced evaluation of text-to-SQL systems, potentially reshaping our understanding of state-of-the-art performance in this field.
title	FLEX: Expert-level False-Less EXecution Metric for Reliable Text-to-SQL Benchmark
topic	Computation and Language Information Retrieval Machine Learning
url	https://arxiv.org/abs/2409.19014

Similar Items