Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lei, Fangyu, Chen, Jixuan, Ye, Yuxiao, Cao, Ruisheng, Shin, Dongchan, Su, Hongjin, Suo, Zhaoqing, Gao, Hongcheng, Hu, Wenjing, Yin, Pengcheng, Zhong, Victor, Xiong, Caiming, Sun, Ruoxi, Liu, Qian, Wang, Sida, Yu, Tao
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910878438260736
author Lei, Fangyu
Chen, Jixuan
Ye, Yuxiao
Cao, Ruisheng
Shin, Dongchan
Su, Hongjin
Suo, Zhaoqing
Gao, Hongcheng
Hu, Wenjing
Yin, Pengcheng
Zhong, Victor
Xiong, Caiming
Sun, Ruoxi
Liu, Qian
Wang, Sida
Yu, Tao
author_facet Lei, Fangyu
Chen, Jixuan
Ye, Yuxiao
Cao, Ruisheng
Shin, Dongchan
Su, Hongjin
Suo, Zhaoqing
Gao, Hongcheng
Hu, Wenjing
Yin, Pengcheng
Zhong, Victor
Xiong, Caiming
Sun, Ruoxi
Liu, Qian
Wang, Sida
Yu, Tao
contents Real-world enterprise text-to-SQL workflows often involve complex cloud or local data across various database systems, multiple SQL queries in various dialects, and diverse operations from data transformation to analytics. We introduce Spider 2.0, an evaluation framework comprising 632 real-world text-to-SQL workflow problems derived from enterprise-level database use cases. The databases in Spider 2.0 are sourced from real data applications, often containing over 1,000 columns and stored in local or cloud database systems such as BigQuery and Snowflake. We show that solving problems in Spider 2.0 frequently requires understanding and searching through database metadata, dialect documentation, and even project-level codebases. This challenge calls for models to interact with complex SQL workflow environments, process extremely long contexts, perform intricate reasoning, and generate multiple SQL queries with diverse operations, often exceeding 100 lines, which goes far beyond traditional text-to-SQL challenges. Our evaluations indicate that based on o1-preview, our code agent framework successfully solves only 21.3% of the tasks, compared with 91.2% on Spider 1.0 and 73.0% on BIRD. Our results on Spider 2.0 show that while language models have demonstrated remarkable performance in code generation -- especially in prior text-to-SQL benchmarks -- they require significant improvement in order to achieve adequate performance for real-world enterprise usage. Progress on Spider 2.0 represents crucial steps towards developing intelligent, autonomous, code agents for real-world enterprise settings. Our code, baseline models, and data are available at https://spider2-sql.github.io
format Preprint
id arxiv_https___arxiv_org_abs_2411_07763
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
Lei, Fangyu
Chen, Jixuan
Ye, Yuxiao
Cao, Ruisheng
Shin, Dongchan
Su, Hongjin
Suo, Zhaoqing
Gao, Hongcheng
Hu, Wenjing
Yin, Pengcheng
Zhong, Victor
Xiong, Caiming
Sun, Ruoxi
Liu, Qian
Wang, Sida
Yu, Tao
Computation and Language
Artificial Intelligence
Databases
Real-world enterprise text-to-SQL workflows often involve complex cloud or local data across various database systems, multiple SQL queries in various dialects, and diverse operations from data transformation to analytics. We introduce Spider 2.0, an evaluation framework comprising 632 real-world text-to-SQL workflow problems derived from enterprise-level database use cases. The databases in Spider 2.0 are sourced from real data applications, often containing over 1,000 columns and stored in local or cloud database systems such as BigQuery and Snowflake. We show that solving problems in Spider 2.0 frequently requires understanding and searching through database metadata, dialect documentation, and even project-level codebases. This challenge calls for models to interact with complex SQL workflow environments, process extremely long contexts, perform intricate reasoning, and generate multiple SQL queries with diverse operations, often exceeding 100 lines, which goes far beyond traditional text-to-SQL challenges. Our evaluations indicate that based on o1-preview, our code agent framework successfully solves only 21.3% of the tasks, compared with 91.2% on Spider 1.0 and 73.0% on BIRD. Our results on Spider 2.0 show that while language models have demonstrated remarkable performance in code generation -- especially in prior text-to-SQL benchmarks -- they require significant improvement in order to achieve adequate performance for real-world enterprise usage. Progress on Spider 2.0 represents crucial steps towards developing intelligent, autonomous, code agents for real-world enterprise settings. Our code, baseline models, and data are available at https://spider2-sql.github.io
title Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
topic Computation and Language
Artificial Intelligence
Databases
url https://arxiv.org/abs/2411.07763