Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Jian, Chen, Zhenyan, Hu, Xuming, Zhou, Peilin, Hua, Yining, Fang, Han, Choy, Cissy Hing Yee, Ke, Xinmei, Luo, Jingfeng, Yuan, Zixuan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2509.14507
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918285273989120
author Chen, Jian
Chen, Zhenyan
Hu, Xuming
Zhou, Peilin
Hua, Yining
Fang, Han
Choy, Cissy Hing Yee
Ke, Xinmei
Luo, Jingfeng
Yuan, Zixuan
author_facet Chen, Jian
Chen, Zhenyan
Hu, Xuming
Zhou, Peilin
Hua, Yining
Fang, Han
Choy, Cissy Hing Yee
Ke, Xinmei
Luo, Jingfeng
Yuan, Zixuan
contents Natural Language to SQL (NL2SQL) provides a new model-centric paradigm that simplifies database access for non-technical users by converting natural language queries into SQL commands. Recent advancements, particularly those integrating Retrieval-Augmented Generation (RAG) and Chain-of-Thought (CoT) reasoning, have made significant strides in enhancing NL2SQL performance. However, challenges such as inaccurate task decomposition and keyword extraction by LLMs remain major bottlenecks, often leading to errors in SQL generation. While existing datasets aim to mitigate these issues by fine-tuning models, they struggle with over-fragmentation of tasks and lack of domain-specific keyword annotations, limiting their effectiveness. To address these limitations, we present DeKeyNLU, a novel dataset which contains 1,500 meticulously annotated QA pairs aimed at refining task decomposition and enhancing keyword extraction precision for the RAG pipeline. Fine-tuned with DeKeyNLU, we propose DeKeySQL, a RAG-based NL2SQL pipeline that employs three distinct modules for user question understanding, entity retrieval, and generation to improve SQL generation accuracy. We benchmarked multiple model configurations within DeKeySQL RAG pipeline. Experimental results demonstrate that fine-tuning with DeKeyNLU significantly improves SQL generation accuracy on both BIRD (62.31% to 69.10%) and Spider (84.2% to 88.7%) dev datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14507
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeKeyNLU: Enhancing Natural Language to SQL Generation through Task Decomposition and Keyword Extraction
Chen, Jian
Chen, Zhenyan
Hu, Xuming
Zhou, Peilin
Hua, Yining
Fang, Han
Choy, Cissy Hing Yee
Ke, Xinmei
Luo, Jingfeng
Yuan, Zixuan
Artificial Intelligence
Computation and Language
Natural Language to SQL (NL2SQL) provides a new model-centric paradigm that simplifies database access for non-technical users by converting natural language queries into SQL commands. Recent advancements, particularly those integrating Retrieval-Augmented Generation (RAG) and Chain-of-Thought (CoT) reasoning, have made significant strides in enhancing NL2SQL performance. However, challenges such as inaccurate task decomposition and keyword extraction by LLMs remain major bottlenecks, often leading to errors in SQL generation. While existing datasets aim to mitigate these issues by fine-tuning models, they struggle with over-fragmentation of tasks and lack of domain-specific keyword annotations, limiting their effectiveness. To address these limitations, we present DeKeyNLU, a novel dataset which contains 1,500 meticulously annotated QA pairs aimed at refining task decomposition and enhancing keyword extraction precision for the RAG pipeline. Fine-tuned with DeKeyNLU, we propose DeKeySQL, a RAG-based NL2SQL pipeline that employs three distinct modules for user question understanding, entity retrieval, and generation to improve SQL generation accuracy. We benchmarked multiple model configurations within DeKeySQL RAG pipeline. Experimental results demonstrate that fine-tuning with DeKeyNLU significantly improves SQL generation accuracy on both BIRD (62.31% to 69.10%) and Spider (84.2% to 88.7%) dev datasets.
title DeKeyNLU: Enhancing Natural Language to SQL Generation through Task Decomposition and Keyword Extraction
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.14507