Evaluating SQL Understanding in Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Rahaman, Ananya, Zheng, Anny, Milani, Mostafa, Chiang, Fei, Pottinger, Rachel
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913545242804224
author Rahaman, Ananya
Zheng, Anny
Milani, Mostafa
Chiang, Fei
Pottinger, Rachel
author_facet Rahaman, Ananya
Zheng, Anny
Milani, Mostafa
Chiang, Fei
Pottinger, Rachel
contents The rise of large language models (LLMs) has significantly impacted various domains, including natural language processing (NLP) and image generation, by making complex computational tasks more accessible. While LLMs demonstrate impressive generative capabilities, there is an ongoing debate about their level of "understanding," particularly in structured domains like SQL. In this paper, we evaluate the extent to which LLMs "understand" SQL by testing them on a series of key SQL tasks. These tasks, such as syntax error detection, missing token identification, query performance prediction, query equivalence checking, and query explanation, assess the models' proficiency in recognition, context awareness, semantics, and coherence, which are essential skills for SQL understanding. We generate labeled datasets from well-known workloads, and evaluate the latest LLMs, focusing on how query complexity and syntactic features influence performance. Our results indicate that while GPT4 excels at tasks requiring recognition and context, all models struggle with deeper semantic understanding and coherence, especially in query equivalence and performance estimation, revealing the limitations of current LLMs in achieving full SQL comprehension.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10680
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating SQL Understanding in Large Language Models
Rahaman, Ananya
Zheng, Anny
Milani, Mostafa
Chiang, Fei
Pottinger, Rachel
Databases
The rise of large language models (LLMs) has significantly impacted various domains, including natural language processing (NLP) and image generation, by making complex computational tasks more accessible. While LLMs demonstrate impressive generative capabilities, there is an ongoing debate about their level of "understanding," particularly in structured domains like SQL. In this paper, we evaluate the extent to which LLMs "understand" SQL by testing them on a series of key SQL tasks. These tasks, such as syntax error detection, missing token identification, query performance prediction, query equivalence checking, and query explanation, assess the models' proficiency in recognition, context awareness, semantics, and coherence, which are essential skills for SQL understanding. We generate labeled datasets from well-known workloads, and evaluate the latest LLMs, focusing on how query complexity and syntactic features influence performance. Our results indicate that while GPT4 excels at tasks requiring recognition and context, all models struggle with deeper semantic understanding and coherence, especially in query equivalence and performance estimation, revealing the limitations of current LLMs in achieving full SQL comprehension.
title Evaluating SQL Understanding in Large Language Models
topic Databases
url https://arxiv.org/abs/2410.10680